American AI labs have complained privately for years that Chinese competitors were farming their models for training data. On September 8, the U.S. government put the accusation on letterhead, with names attached.
Joint cybersecurity advisory AA26-251A, from the National Security Agency, the Cybersecurity and Infrastructure Security Agency and the FBI, alleges that six China-based companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI — have run sustained knowledge distillation campaigns against U.S. frontier models since at least late 2024, extracting billions of tokens across millions of exchanges. The named targets: Anthropic's Claude, OpenAI's GPT, Google's Gemini and xAI's Grok.
Distillation, the agencies concede, is a legitimate machine learning technique: train a smaller student model on a larger teacher's outputs. What they allege here is something else — systematic extraction of proprietary functionalities and capabilities through campaigns that “form the core—not merely a supplement—of their AI development strategy.” The activity took place, the agencies write, “likely with Chinese government awareness.”
What the Advisory Alleges
The detail is unusually granular for a public advisory. DeepSeek, according to the agencies, distilled from a dozen models including Claude Opus 4.1, Gemini 2.5 Pro Preview, GPT-5 and Grok 4 between late 2024 and mid-2025 to generate synthetic training data for its R1 and V3 models. The advisory adds a pointed economic claim: DeepSeek's widely quoted $5.6 million training cost is misleading, the agencies say, because it excludes the cost of data acquired through distillation.
Moonshot AI is alleged to have distilled 18 different U.S. models, including Anthropic's current flagship, to train Kimi-K2 and Kimi-K3. Alibaba allegedly distilled Claude and GPT-5 outputs into its Qwen family. MiniMax is accused of deploying prompt injections meant to convince Claude Code it was a MiniMax product, and of redirecting extraction traffic to a new Claude model within 24 hours of release. Z.AI, the agencies say, had distilled billions of tokens of GPT-5.5 and Claude Opus 4.8 data by mid-2026.
The tradecraft section reads like a fraud investigation: requests routed through cloud providers and aggregators to obfuscate metadata; a gray market of API proxies known as “transfer stations” used to bypass geographic restrictions; bulk-purchased premium subscriptions shared across teams; automated failover when a pathway is blocked; and evaluation frameworks built to detect whether a provider has started degrading responses. Anthropic reported in February that three Chinese labs had generated more than 16 million exchanges with Claude through roughly 24,000 fraudulent accounts.
The recommended mitigations are three: detect anomalous prompts, accounts and networks; deploy targeted response changes that subtly degrade or alter answers to suspected distillation traffic, varying them across requests so the distiller cannot benchmark its way around them; and build intelligence sharing among model providers, cloud platforms and API aggregators. Notably, the advisory recommends against telling suspected distillers that their responses have been altered, while urging that safety researchers and third-party evaluators be informed. “We strongly urge AI companies to take immediate steps to safeguard their platforms against knowledge distillation campaigns,” said CISA Acting Director Nick Andersen.
Beijing rejected the allegations within a day. China's Ministry of Commerce called them “unfounded both in fact and in law,” accusing Washington of “politicizing and weaponizing distillation, which is a normal technical and commercial practice within the industry.” Foreign Ministry spokesperson Mao Ning urged the U.S. to “refrain from making unfounded accusations against China.” None of the six named companies had issued a public response as of publication; Moonshot AI declined to comment to NPR on a similar White House accusation in July, though Chinese media reported it denied the claim.
What the Advisory Can and Cannot Do
Read closely, AA26-251A is careful in a way the headlines around it are not. It never uses the words illegal, unlawful, theft, crime, copyright or trade secret. What it alleges is violation of terms of use — a contractual matter, not a criminal one. The hedges do real work: “likely with Chinese government awareness” is the extent of the attribution to Beijing, and the underlying evidence is not disclosed.
That gap matters, because the legal ground is unsettled. Bahrad Sokhansanj of the Institute for Law and AI, writing in Lawfare in July, argued that distillation is not model theft in any recognized sense: copyright is a poor fit because AI outputs generally are not copyrightable, and trade secret protection is hard to claim over outputs any paying customer can obtain. “But if every TOS violation counts as ‘theft,'” he wrote, “then the concept has no limiting principle.” The stronger theory runs through the Computer Fraud and Abuse Act — not for the distillation itself, but for the fraudulent accounts and credential misuse used to evade access cutoffs. After Van Buren v. United States, ordinary terms-of-service violations do not exceed authorized access.
There is also an awkwardness the advisory does not address. Georgetown Law's Anupam Chander has noted the irony: “the AI companies believe, and they have argued, that learning from others is a perfectly fair use of other people's copyrighted work.” Elon Musk testified at the Musk v. Altman trial that “generally AI companies distill other AI companies.” Distillation from open-weight models — many of them Chinese — is standard practice in U.S. labs, a point Beijing made explicitly. An enforcement regime that treats output-learning as property could constrain the open ecosystem far more than six well-resourced firms operating through proxies.
Proving distillation is its own problem. The forensic signal is statistical, not a smoking gun, and any provider acting on those heuristics risks throttling legitimate heavy users.
Watch three things. Whether Treasury follows through on the sanctions and Entity List designations Secretary Scott Bessent floated in July. Whether the Deterring American AI Model Theft Act, advanced by the House Foreign Affairs Committee in April, moves in a form that creates new rights in model outputs. And whether the promised U.S.-China intergovernmental AI dialogue survives an advisory Beijing has already called an act of technological hegemony.
“But if every TOS violation counts as theft, then the concept has no limiting principle.”— Bahrad A. Sokhansanj, Senior Research Scholar, Institute for Law and AI