TL;DR
- On September 8, the NSA, FBI, and CISA jointly published advisory AA26-251A naming six Chinese AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI — as running "industrial-scale" distillation campaigns against US frontier AI models since late 2024. - The agencies say the six companies extracted "billions of tokens across millions of exchanges" from variants of Claude, GPT, Gemini, and Grok, and that the activity was "likely conducted with Chinese government awareness." - Anthropic separately disclosed that the three named labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts. It later told lawmakers that Alibaba's Qwen lab alone ran over 28.8 million exchanges through nearly 25,000 fraudulent accounts — its largest such disclosure. - The NSA/FBI/CISA advisory says DeepSeek's widely quoted $5.6 million training cost for R1 is "misleading" because it excludes the cost of distilled data acquired through the campaigns. - Treasury Secretary Scott Bessent has threatened sanctions and Entity List designations. - The most unusual element of the advisory: agencies recommend that US labs quietly degrade responses to high-confidence distillation requests without informing the users making them — turning API responses into an active countermeasure surface. - China's Ministry of Commerce calls the accusations groundless and warns of "resolute countermeasures" if the US suppresses Chinese AI companies under the pretext of targeting distillation.
Advisory AA26-251A, published September 8, 2026, is the most detailed public accounting of alleged AI intellectual property theft ever published by US security agencies. It is also the document that reframes a technical practice — model distillation — as a nation-state intelligence operation.
Here is what it actually says, what it doesn't say, and what it means for enterprise teams currently using or evaluating Chinese AI models.
What distillation is and what makes this different
Distillation is a standard technique in machine learning. A smaller "student" model is trained to mimic the outputs of a larger, more capable "teacher" model. The student learns not just the right answers but how the teacher reasons under uncertainty. Done with permission, it is ordinary practice at every major AI lab, American ones included. American labs have themselves built models using Chinese open-source models as inputs.
The advisory does not dispute that distillation is legitimate. It draws the line at three things: training on outputs obtained in breach of terms of service, deliberately targeting a competitor's proprietary capabilities at coordinated scale, and actively working to avoid detection.
The evasion architecture described in the advisory is what pushes this from a trade complaint toward a threat report. The six named companies are accused of routing queries through "transfer stations" — a gray market of proxy services marketed on Chinese platforms — to bypass geographic restrictions, evade safeguards, and eliminate traceability. Accounts were distributed across dozens of thousands of fraudulent profiles. Premium subscriptions were purchased in bulk and shared across development teams to make industrial-volume extraction look like ordinary paid usage. Automated failover between platforms triggered when any single access channel was blocked.
This is not a researcher at a Chinese lab running a few thousand comparison queries. The advisory describes a coordinated, multi-year, multi-company operation with evasion infrastructure built specifically to resist detection.
What was taken and who took it
The company-specific allegations are specific.
DeepSeek is accused of running organized distillation campaigns since late 2024, targeting reasoning capabilities, coding performance, and domain-specific functions. The advisory names four versions of Claude, two versions of Gemini, five versions of ChatGPT, and Grok 4 as sources. DeepSeek's R1 and V3 models are named as beneficiaries. The agencies call DeepSeek's $5.6 million training cost "misleading" because it excludes the value of data obtained through the campaigns.
Moonshot AI is accused of extracting Claude Fable 5 data to train Kimi-K3, and GPT-4o data to train Kimi-K2. The advisory says Moonshot distilled 18 different US models across its two flagship systems.
Alibaba is accused of distilling Claude-4, Claude Opus, Claude Sonnet, and GPT-5 to improve software engineering, customer service, and image and character creation in its Qwen family of models. Anthropic told lawmakers that operators linked to Alibaba's Qwen lab ran over 28.8 million exchanges through nearly 25,000 fraudulent accounts — the largest distillation campaign Anthropic has publicly disclosed.
MiniMax is named in connection with its M2 model, accused of extracting chain-of-thought reasoning, reinforcement learning, and supervised fine-tuning capabilities from Claude Code, Claude Sonnet 4, Claude Opus, and multiple Gemini versions. MiniMax is also accused of using prompt injection to make Claude Code believe it was a MiniMax product during internal software development.
StepFun is named in connection with its Step 4 model. The same day as this article, StepFun launched Step 5 Preview — its new flagship agentic model.
Z.AI is accused of distilling billions of tokens from GPT-5.5 and Claude Opus 4.8 to develop chain-of-thought reasoning capabilities by mid-2026.
Google's Threat Intelligence Group separately reported this week that some distillation campaigns against its Gemini models exceeded 100 million prompts, targeting visual understanding, audio understanding, image generation, and video generation.
The unusual remedy
The most striking element of advisory AA26-251A is not the accusations. It is the recommended countermeasure.
Rather than advising American labs to simply block suspect traffic — which the advisory notes is detected and bypassed through automated failover — the agencies recommend that labs quietly degrade responses to high-confidence distillation requests without informing users. The degradation should be varied and unpredictable, so it cannot be measured and accounted for by the extraction operation.
This is a significant policy position. The agencies are recommending that US AI companies deceive users they believe are running distillation campaigns — delivering intentionally degraded outputs without disclosure. For enterprise customers, this has a direct implication: if you are using a US frontier model's API and your usage patterns resemble high-volume coordinated extraction, your responses may already be degraded without your knowledge.
The advisory does not define the threshold for "high-confidence distillation request." The monitoring signatures it describes — highly coordinated prompt texts submitted at volumes in the thousands to millions on narrow topics, prompts explicitly requesting chain-of-thought externalization, sustained multi-month campaigns targeting a single technical domain — are also behaviors that large-scale enterprise AI deployments can exhibit legitimately.
What China said
China's Ministry of Commerce called the accusations groundless, without legal basis, and "the latest manifestation of technological hegemony and computing power monopoly." A ministry spokesperson argued that distillation is a neutral technique used by developers worldwide, including American ones. The ministry pointed out that China's open-source models are available to global firms including US companies, some of which have disclosed using Chinese models as inputs.
The ministry warned that "if the U.S. suppresses Chinese AI companies under the pretext of targeting distillation, China will take resolute countermeasures."
This framing is not inaccurate about the technique itself. Distillation is neutral. The dispute is about authorization, scale, and evasion architecture. China's public position does not engage with the specific allegations about fraudulent accounts, transfer station infrastructure, or the campaigns' volume and duration.
The enterprise implication
Enterprise teams currently using DeepSeek R1, Alibaba Qwen, Kimi K3, or other models from the named labs are operating in supply chains that the US government has now formally characterized as national security concerns.
This does not mean those models are technically inferior or unsafe for all use cases. It does mean three things.
First, the advisory's concern about safety guardrails applies directly. Anthropic has previously stated that models built through illicit distillation are unlikely to retain the safety safeguards of the models they were trained on — meaning the safety profile of the source model does not transfer to the distilled derivative. Enterprise AI governance frameworks that rely on the safety certifications of US frontier models should not assume those certifications extend to derivatives built through distillation campaigns.
Second, any enterprise operating in a regulated industry or under US government contracts should involve legal and compliance review in vendor selection decisions that include the six named companies. Treasury Secretary Bessent has explicitly named sanctions and Entity List designations as potential consequences. Entity List designation would restrict US companies from doing business with the named firms.
Third, the degraded-response countermeasure the advisory recommends creates a new due diligence question for enterprise AI procurement. If your vendor is following the advisory's guidance, the model you evaluated in a benchmark may not be the model your agents are running against in production under certain query patterns.
The timing
Advisory AA26-251A was published September 8, two weeks before Trump's state dinner for Xi Jinping on September 24. Sam Altman, Jensen Huang, and Tim Cook are on the guest list, with AI explicitly on the diplomatic agenda. The advisory — the most aggressive formal accusation of AI IP theft the US government has published — lands in the same ten-day window as the three-lab safety coordination announcement, the kill switch debate, and Amodei's slowdown call.
The US government published a document calling six Chinese AI companies state-assisted IP thieves and recommending that American labs deceive their users to stop them — the same week the US president invited the leaders of American AI to dinner with the Chinese head of state.
The policy environment around AI is moving faster than most enterprise governance frameworks are tracking. That gap is the risk.


