Summary of Key Points
On August 6th, two significant events in the AI community collided: the founder of Google DeepMind resigned from his role as CEO due to the underperformance of their Gemini model; Zhang Yiming from ByteDance made the decision that "large models should not be distilled." This decision has sparked a divide within the industry—some see it as a publicity stunt, while others praise it for returning to more fundamental principles. The article discusses the nature of "model distillation," the underlying reasons behind ByteDance's decision, the controversy within the industry, and the diverse paths taken by Chinese AI companies. It argues that although ByteDance has chosen a "costly and difficult" path, this diversity is precisely what the Chinese AI ecosystem needs.
Detailed Analysis
1. Understanding What Model Distillation Is—and Why Everyone Uses It
Model distillation is not a chemical process but a method in AI training, which can be categorized into two types:
- Using Your Own Model: Training a smaller model with your own large model (for example, using GPT-4 to train GPT-3.5), which is a common and uncontroversial practice in the industry.
- Using Someone Else's Model: Using the output of another model (such as ChatGPT) to train your own model, which is the subject of much debate.
To illustrate, using someone else’s model is like a student copying a top student’s homework—it can quickly improve performance (in model rankings), but it may not truly teach the underlying reasoning skills. Why do people use this method? Because it’s effective and saves computational resources. For instance, DeepSeek used data from 800,000 R1 models to significantly improve the performance of a smaller model (AIME24 benchmark) without incurring high computational costs.
However, problems have arisen: the United States has begun to use model distillation as a reason for sanctions. The White House has accused some Chinese companies of stealing technology, and Anthropic claims that these companies use fake accounts to collect data, even identifying distillation activities through IP and access patterns.
2. Zhang Yiming’s Decision to “Never Distill” Models: Fear of Punishment or a Firm Standpoint?
Many believe ByteDance is avoiding issues related to TikTok’s dealings with the US government. However, the article provides counterarguments: they even prohibit the use of domestic open-source models (like Kimi K3), which have loose licensing and pose no risk to TikTok. So, what are the real reasons behind this decision?
- Technical Purism: Last year, ByteDance released an open-source model that did not contain synthetic data, fearing that such data could contaminate their research foundation.
- TikTok’s Global Presence: Although not the primary reason, ByteDance has the largest overseas business and does not want to be held accountable for any potential issues (recall how Project Seed was suspended by OpenAI for suspected model distillation).
- Zhang Yiming’s Personal Commitment: After stepping down as CEO, Zhang focused on AI research, attending team reviews, staying up late to read papers, and hiring a professor from Singapore for additional guidance. His decision to avoid model distillation reflects his commitment to developing genuine AI capabilities.
3. Is Model Distillation a “Shortcut” or a “Dead End?”
The debate within the industry is intense:
- Proponents of Distillation: It’s effective and saves resources. Research from Tsinghua University shows that distilled models outperform base models, and even Anthropic acknowledges it as a legitimate training method.
- Critics of Distillation: Repeatedly using the same or others’ outputs can lead to information loss (similar to how a message loses meaning after being passed around multiple times), a phenomenon known as “model collapse.”
- Neutral Viewpoint: AI researcher Lambert argues that distillation is akin to imitation, while reinforcement learning represents true exploration—developing unique capabilities through trial and error. He also notes that using synthetic data wisely (with various sources and real-world data) can prevent model collapse.
4. Chinese AI Companies: More Than Just Model Distillation
The article provides three examples showing that distillation is just one tool among many:
- Zhipu GLM-5.3: This model improved its capabilities through reinforcement learning and a long-context architecture, without relying on distillation or additional parameters. The developers claim that the potential of this model’s intelligence is far from being fully explored.
- DeepSeek: This company trained its models through self-play, independent of external outputs, and even OpenAI’s chief researcher acknowledged their independent discovery of core concepts related to model performance.
- Alibaba’s Qianwen: Although accused of distillation, it serves as a foundational model for many smaller models worldwide, demonstrating its robust capabilities.
In conclusion, the debate over model distillation is largely about stigmatizing Chinese AI. Distillation may affect development speed but does not limit potential.
5. ByteDance’s Bet: Can They Win by Taking the Difficult Path?
ByteDance is exploring the possibility of training models with over 5 trillion or even 10 trillion parameters. What are the challenges?
- Short-term Costs: Liang Rubo has admitted that the gap with leading overseas models is widening.
- Long-term Value: If successful, ByteDance could become a major player in the global AI landscape and rewrite the history of Chinese AI development. Failure would result in a costly lesson for the industry.
However, this diversity of approaches is beneficial, as different Chinese companies need to explore unique paths. Zhang Yiming’s decision deserves respect, but it shouldn’t be imposed as an industry standard.
Final Thought
Model distillation is not about taking a stance; it’s a technical tool. ByteDance has chosen the toughest path, and their courage to go against the norm is essential for AI innovation. Regardless of the outcome, this approach adds another possibility to China’s AI ecosystem.