US AI Labs Accuse Chinese Rivals of Using Model Distillation to Steal Capabilities
Washington and leading AI firms warn that 'student' models are being trained on proprietary data to bypass massive R&D costs.
Washington and leading U.S. AI firms are sounding the alarm over 'model distillation,' a technique they claim Chinese rivals are using to systematically extract capabilities from proprietary American AI models. The dispute marks a significant escalation in the tech rivalry, as U.S. companies allege that Chinese entities are harvesting advanced reasoning and software engineering skills without consent.
At the center of the conflict is the process of model distillation, where a large, powerful 'teacher' model generates examples—such as complex answers or code—that are then used to train a smaller 'student' model. This allows the smaller model to perform specific tasks with high efficiency without requiring the massive computing power or vast datasets typically needed to build a frontier model from scratch. Anthropic has specifically accused Chinese entities, including DeepSeek, Moonshot, and MiniMax, of conducting large-scale campaigns to obtain capabilities from its Claude models.
The Shift to Reasoning Traces
Distillation is not a new or inherently illicit practice; it is a standard AI research tool used globally. U.S. institutions have employed the technique for years, including Microsoft through its Orca research and Stanford University via the Alpaca project. However, the current tension stems from the use of 'closed' proprietary models versus 'open-weight' models.
The battleground has recently shifted toward 'reasoning traces'—the step-by-step logic a model uses to reach a conclusion. These traces are far more valuable for training smaller models than final answers alone. Florian Tramèr, an assistant professor at ETH Zurich, illustrated the difference by noting that a student has a much harder time learning from a book of math problems with only final solutions than from one providing detailed, step-by-step instructions.
Implications for the AI Arms Race
This conflict represents a new front in the global AI arms race. If Chinese firms can successfully distill frontier capabilities into smaller, cheaper models, it directly undermines the competitive advantage and the multi-billion dollar R&D investments made by U.S. AI labs.
Furthermore, this practice complicates U.S. efforts to maintain a technological lead through chip export controls. While Washington has restricted the flow of high-end GPUs to China to slow AI development, distillation provides a workaround. By leveraging the outputs of existing American models, Chinese developers can create high-performance AI using significantly fewer high-end chips.
What to Watch
As the dispute intensifies, the industry is watching whether U.S. firms will implement stricter technical barriers to prevent distillation or if the U.S. government will introduce new regulatory frameworks to protect proprietary model outputs. It remains to be seen how Chinese firms will respond to these accusations or if they will move toward more transparent development paths to avoid further diplomatic and commercial friction.