The U.S. wants to contain China's AI. Silicon Valley keeps using it.
Key Takeaways
- ▸Model distillation is a standard technique used by all major AI laboratories worldwide, not a uniquely Chinese practice or security threat
- ▸Silicon Valley companies actively incorporate advances from China's open-weight models while the U.S. government simultaneously restricts Chinese access to American frontier models, revealing a policy contradiction
- ▸As frontier models become more capable, their outputs increase in value as training material, intensifying competition over whose models become the industry standard teachers
Summary
Anthropic's chief national security officer and former Biden administration export control architect Tarun Chhabra recently warned that Chinese AI developers like Zhipu are using model distillation to extract capabilities from frontier American models. However, the narrative became complicated almost immediately when Mira Murati's Thinking Machines revealed that its first foundation model incorporated components from Chinese models including DeepSeek-V3 and Moonshot AI's Kimi K2.5.
The contradiction highlights a fundamental tension in U.S. AI policy. While American firms are tightening API protections and implementing behavioral monitoring to prevent unauthorized extraction of their frontier capabilities, they are simultaneously drawing from China's rapidly improving open-weight model ecosystem. Model distillation—where a larger teacher model generates outputs used to train a smaller student model—is not new technology, having been part of machine learning for over a decade.
The reality is that distillation has expanded far beyond simple model compression. Modern distillation now encompasses synthetic data generation, reasoning improvements, multilingual performance, and alignment training. Nearly every major AI laboratory—including OpenAI, Anthropic, Google DeepMind, Meta, Alibaba, Tencent, Moonshot, DeepSeek, and Zhipu—employs some version of teacher-student training. The competitive question has shifted from whether laboratories use distillation to whose models become the teachers, as the value of model outputs has grown dramatically with frontier model capabilities.
Editorial Opinion
The juxtaposition of Anthropic warning about Chinese distillation while Thinking Machines admits to using Chinese models exposes fundamental incoherence in current U.S. AI policy. As model distillation becomes a standard practice deployed by every major laboratory worldwide, policymakers face a hard constraint: restricting access to American frontier models while Silicon Valley actively incorporates advances from Chinese open-weight systems. Sustainable AI governance may require acknowledging that competitive development inherently involves learning from previous innovations.


