U.S. Agencies Warn China Used AI Distillation to Copy Frontier Models
CISA, NSA, and FBI detail how Chinese companies extracted billions of tokens from Claude, GPT, Gemini, and Grok through industrial-scale knowledge distillation campaigns.
U.S. Agencies Warn China Used AI Distillation to Copy Frontier Models
CISA, the NSA, and the FBI have issued a joint cybersecurity advisory detailing how Chinese AI companies conducted large-scale knowledge distillation campaigns to replicate capabilities from leading U.S. AI models. The agencies report that firms including DeepSeek, Moonshot AI, Alibaba Group, MiniMax, StepFun, and Z.AI extracted billions of tokens through millions of requests targeting Claude, GPT, Gemini, and Grok variants since at least late 2024.
According to the advisory, first reported by Help Net Security, the activity likely occurred with the knowledge of the Chinese government. CISA Acting Director Nick Andersen urged AI companies to take immediate protective measures against distillation campaigns that threaten to erode the competitive advantage of American firms.
Why it matters
This advisory represents the first detailed public disclosure by U.S. intelligence and cybersecurity agencies of coordinated AI capability theft at industrial scale. The findings challenge public claims about training costs—the agencies note DeepSeek's reported $5.6 million training figure for its models doesn't account for the true cost of capabilities obtained through extensive distillation from U.S. systems. For AI providers, the disclosure signals that nation-state actors are systematically exploiting API access to bypass the computational costs and technical challenges of developing frontier capabilities independently.
How the campaigns operated
The Chinese companies employed sophisticated techniques to evade detection and restrictions. They used a gray market of API proxies called "transfer stations" to bypass regional restrictions and obscure their country of origin. Pools of premium accounts distributed requests across systems to avoid usage limits, while centralized routing systems directed traffic across APIs, cloud providers, and aggregators based on availability and quotas.
DeepSeek used extracted data and capabilities to train its R1 and V3 models, focusing on reasoning, writing, agentic functions, and specialized tasks including legal work. Moonshot AI targeted software engineering, mathematics, and reinforcement learning capabilities for its Kimi models starting in mid-2025. By mid-2026, Z.AI had distilled billions of tokens from GPT-5.5 and Claude Opus 4.8 to develop chain-of-thought reasoning capabilities.
Detection and mitigation strategies
The agencies outlined specific indicators that providers can monitor, including continuous 24/7 usage without normal human patterns, anomalous subscription-to-API usage ratios, new subscriptions that rapidly reach maximum usage, and coordinated behavior across account pools.
Recommended defenses include stronger account identity verification, varying responses to suspected distillation attempts, limiting reasoning depth in outputs, and routing suspected operators to less capable models without notification. For confirmed malicious activity, providers can alter responses while informing AI safety researchers and third-party evaluators of model changes.
Differential privacy—adding controlled noise to model outputs—offers another defense layer, though stronger privacy protections can reduce model accuracy. The agencies recommend combining differential privacy with rate limits, response controls, and monitoring. Sharing indicators such as IP addresses, domains, query volumes, and timing patterns across providers can help identify coordinated campaigns that might otherwise appear isolated.
The advisory maps the distillation campaigns to the MITRE ATLAS framework, covering tactics from resource development through AI model access, execution, defense evasion, and exfiltration. Details were first reported by Help Net Security.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

