第一财经

"The Distillation Debate: A Fight for Control over the Development of the Intelligent Era" | Expert Opinions

原文:蒸馏之争,争的是智能时代的发展权|专家热评

The US Authorities Accuse China of “Malicious Distillation” in AI? Don’t Be Misled—It’s Really About Who Has the Upper Hand

Hello everyone, I’m your financial journalist and economist. Recently, the tech world has been in an uproar: the National Security Agency (NSA), the Cybersecurity and Infrastructure Security Agency (CISA), and the Federal Bureau of Investigation (FBI) jointly issued a statement accusing Chinese AI companies of engaging in “industrial-scale” “malicious” distillation, which essentially amounts to accusing us of stealing their technology.

This sounds quite alarming, as if we’ve committed some serious crime. But if you break down the logic behind it, you’ll see that this isn’t about a technical dispute at all; it’s a carefully crafted political stunt. Today, I’ll use plain language to揭 the curtain on what’s really going on.

Summary: A Game of Double Standards

In one sentence: The US has politicized a neutral technology called “knowledge distillation” and, using complaints from certain companies, is trying to use state power to maintain its monopoly in AI computing power and suppress the development of Chinese AI.

The core logic is as follows:

1. Technology is innocent: Distillation is a common practice in the AI field; it’s like students using a teacher’s textbook, not stealing.

2. The accusations lack solid evidence: The US’s evidence mainly comes from self-narrations by American companies like Anthropic, without any concrete technical proof, making the argument flawed.

3. Motives are ulterior: The US bans Chinese use of chips while accusing them of learning technology, all in an attempt to create a “data moat” to prevent China from catching up.

4. Double standards: American companies widely use Chinese open-source models (such as DeepSeek and Qwen2) for their own innovation but then accuse Chinese companies of learning from American models.

5. The ultimate goal: To maintain technological hegemony, turning commercial competition into a national security issue to justify further sanctions.

---

Deep Dive: Understanding the Controversy from Five Perspectives

1. What is “distillation,” and is it really “stealing”?

First, let’s clarify what distillation is. In the AI community, knowledge distillation is a well-established and mature technique.

  • Simple analogy: Imagine an experienced professor (a large model) condensing their life’s work into a concise textbook (a smaller model). Students (small models) can quickly grasp the core knowledge without having to read all the books from scratch.
  • Technical definition: It’s about using a large model to train a smaller, faster, and more energy-efficient model for use on devices like phones and tablets.
  • Historical context: This technique was proposed by Nobel laureate Geoffrey Hinton in 2015. OpenAI itself provided distillation services, and Alibaba quickly followed with its DistilQwen2 series after launching Qwen2.
  • Industry consensus: The US industry is well aware of this. NVIDIA CEO Jensen Huang, Microsoft, Meta, and 25 other organizations jointly stated that distillation is a common method of model improvement and should not be confused with illegal theft.

Conclusion: Distillation is a tool for accelerating technological progress and efficient knowledge dissemination. Calling it “malicious stealing” is as absurd as saying students copying notes is stealing.

2. Is the US’s evidence solid? Actually, it’s quite weak

The joint statement by the three US agencies sounds dramatic, but upon closer inspection, the evidence is flimsy, almost self-contradictory.

  • Single-source evidence: The main evidence comes from self-narrations by American companies like Anthropic. It’s like the neighbor accusing you of stealing based on their word alone—would the police act on that?
  • Lack of solid proof: Anthropic claims that tens of thousands of accounts accessed their models, but they can’t provide direct evidence that this data was used to train Chinese models. They rely on indirect clues like IP addresses and metadata.
  • Issue 1: IP addresses can be masked by proxy servers and resellers, making it impossible to accurately identify the companies.
  • Issue 2: Just because you access data doesn’t mean you use it for training. You might just be testing or browsing.
  • Technical black box: Large models are “black boxes,” and it’s difficult to determine the specific data used from the final model weights. It’s like claiming to know the origin of each raisin in a cake.
  • Independent tests: Chinese research found that American models (e.g., different versions of Claude) are more similar to each other, while Chinese models are not significantly different from American ones. This suggests that similar answers are expected in such scenarios, ruling out plagiarism.

Conclusion: The US’s accusations lack solid technical evidence and are based on speculation and the commercial concerns of certain companies.

3. American companies’ double standards: Using Chinese models while accusing China

This is the most ironic point. The US accuses China of malicious distillation, yet American companies are also extensively using Chinese models.

  • Examples:
  • NVIDIA: Used millions of data generated by DeepSeek-R1 to train its Nemotron models and made this public.
  • Startup Perplexity: Developed its own R1-1776 model based on DeepSeek-R1.
  • Hugging Face: Launched Open-R1, which replicates DeepSeek-R1.
  • Microsoft and Amazon Web Services: Have integrated Chinese models like Qwen2 into their platforms.
  • Data comparison: There are over 100,000 derivatives of Qwen, with massive downloads, while the usage of American models in China is much lower than that of Chinese models in the US.

Conclusion: American companies benefit from Chinese open-source models and yet accuse Chinese companies of learning from them. This double standard reveals their true intention to maintain their monopoly.

4. The real purpose: From chip bans to data bans

Why are the US authorities making such a fuss now? Because simply banning chips is no longer enough; they want to build an additional barrier.

  • Extension of computing power monopoly: The US previously used chips (e.g., NVIDIA GPUs) to restrict China’s ability to train large models. Now, they realize that China can still develop high-performance models despite limited computing power through algorithm optimization and open-source ecosystems.
  • Building a data moat: If model outputs are considered “protected assets,” any use of these outputs could be deemed infringement. This way, only the largest US companies with the most data (like OpenAI and Anthropic) can safely innovate, while others face significant legal risks.
  • Politicized commercial dispute: The timeline is clear: OpenAI pressured Congress in February → White House issued a memo in April → Request for enhanced cooperation in June → Joint statement in September. This is a clear path of “company pressure → government endorsement → regulation.” They’re turning a commercial competition (which company has better technology) into a national security issue to justify further sanctions, investment restrictions, and even military intervention.

Conclusion: The goal is not to protect intellectual property but to maintain market share and technological hegemony. They’re afraid of competitors closing the gap through legitimate means.

5. The fate of such bans: They’re doomed to fail

One of the main arguments is that Chinese users violated the terms of service (ToS) of American companies like Anthropic, such as geographical restrictions.

  • Aggressive terms: A US company unilaterally bans Chinese entities from using its services and then accuses users who bypass these restrictions of breaking the rules. It’s like a restaurant posting a sign saying “No foreigners allowed” and then arresting those who climb over the wall to eat.
  • Lack of industry consensus: There’s no international rule prohibiting distillation. Meta’s Llama model allows the use of its outputs for training other models, and NVIDIA encourages the use of synthetic data. This shows that banning distillation is a commercial strategy of certain US companies, not an industry standard.
  • Historical lessons: Measures like DVD region codes and music DRM have been circumvented by the market and technology. Technological bans only delay, not stop the spread of knowledge.

Conclusion: Trying to monopolize technology through contracts and state power goes against the laws of technological development and the trend of global cooperation.

---

For Everyone

1. Stay rational and don’t believe rumors: Don’t panic or get angry at headlines like “US accuses China of stealing AI technology.” Understand the political context and the technical reality.

2. Support open source: The rise of Chinese AI is largely due to open-source ecosystems like Qwen, DeepSeek, and Llama. Open source allows global developers to innovate, which is the strongest weapon against monopolies.

3. Focus on autonomy and control: Although distillation is a legitimate technique, autonomy in core computing power (chips) and foundational frameworks is crucial. National investment and corporate R&D are key for long-term competition.

4. Hope for dialogue and cooperation: AI’s risks (e.g., deepfakes and biosecurity) are global, requiring joint governance by China, the US, and the world. Turning AI into a geopolitical weapon harms the well-being of all.

In conclusion: The US authorities’ statement is a carefully orchestrated political stunt that masks their own anxieties about AI innovation and reveals their ambition to maintain technological hegemony. For China, the best response is to continue to promote open source, enhance core technical capabilities, and use stronger products and broader international cooperation to speak through facts.

Technology knows no borders, but the market does have rules. Those who respect rules, innovation, and cooperation will win the future.