虎嗅

"Before the 10-trillion-parameter mega-model even emerges, ByteDance has already established a new department at the corporate level."

原文:10万亿参数大模型还没影,字节先造了个一级部门

Summary of Key Points

ByteDance has recently integrated the data teams scattered across its various business units into a primary department called “AI Data and Security.” This move reflects the company's ambition to develop super-large-scale AI models with over 10 trillion parameters. Since founder Zhang Yiming firmly opposes the practice of “distilling external models” (i.e., using pre-trained models from others), ByteDance must solve its own data supply issues. The entire large-model industry is currently facing a shortage of high-quality data: the valuable text on public internet platforms is almost depleted, forcing companies to compete for data through methods such as purchasing licenses, hiring experts, and even scanning physical books, often leading to copyright disputes (for example, Anthropic paid $1.5 billion in damages for copyright infringement). ByteDance’s organizational change is essentially an effort to prepare in advance for the challenges of developing these super-large models.

Detailed Analysis

1. Upgrading the Data Department to a Primary Level: Not a Casual Decision, but a Necessity

The newly established “AI Data and Security” department is not created from scratch; it consolidates several previously separate teams, including Global Data, DMC (Data Middle Platform), and Flow’s AIDP platform. Why make this move? Large models are like “super students”—the more parameters they have (equivalent to a more developed brain), the more and higher-quality “teaching materials” (data) they require. ByteDance aims to create a model with 10 trillion parameters, which is several times larger than typical models, meaning the amount of training data needed will also increase significantly. The scattered teams were inefficient due to duplicate efforts; by integrating them into a single primary department on par with Seed (the large-model division) and TikTok, the company is treating data as a core project. After all, without sufficient and high-quality data, even the most powerful computing power and algorithms are useless.

ByteDance has already invested heavily in data: its Seedance data evaluation team consists of thousands of people (one algorithm engineer for every ten data analysts), and this year’s budget for data-related activities in the areas of world models and coding has reached tens of millions of dollars. This is no longer about outsourcing data labeling; it involves a comprehensive process from procurement, cleaning, to synthesis, and evaluation.

2. Zhang Yiming’s Opposition to “Copying Others’ Work”: A Short-term Sacrifice for Long-term Success

In August, it was reported that Zhang Yiming explicitly opposed the use of pre-trained models at a meeting with Seed’s entire team. “Distilling” external models means compressing someone else’s trained results, which can help catch up quickly in the short term but leads to long-term dependence on others and a lack of core competitiveness. By choosing this approach, ByteDance is sacrificing immediate benefits to build its own data foundation. It’s like students who must study from textbooks and do practice exercises rather than copy their classmates’ work—although slower, it ensures a solid foundation. The upgrade of the data department supports this long-term strategy, emphasizing the need for a dedicated team to handle all aspects of data acquisition and management.

3. The Rapid Consumption of Data by Large Models: Public Internet Sources Are Insufficient

The logic behind large-model development used to be simple: more parameters, stronger computing power, and more data. However, this is no longer the case due to a shortage of data. A paper by DeepMind (Chinchilla) states that doubling the number of parameters requires doubling the training data volume for optimal performance. The amount of high-quality human-generated text on public internet platforms (such as news, books, and papers) is limited. Research firm Epoch AI estimates that global public text amounts to about 300 trillion tokens, which could be exhausted between 2026 and 2032 at the current rate of consumption. If models need to undergo “overtraining” (repeatedly using the same data to reduce inference costs), this timeline may even be earlier.

As a result, companies are seeking data from other sources, such as paid news services, specialized databases, and even undigitized physical books.

4. Companies Going All Out to Secure Data: Buying Licenses, Hiring Experts, and Even Scanning Physical Books

The data shortage has prompted companies to take drastic measures:

  • Purchasing Licenses: Google pays Reddit $60 million annually for forum content; OpenAI collaborates with News Corp. to use content from The Wall Street Journal and The Times.
  • Hiring Experts: Companies hire engineers and doctors in fields like coding, mathematics, and science to create and evaluate data, as ordinary people cannot produce high-quality professional content.
  • Scanning Physical Books: Anthropic launched “Project Panama,” purchasing millions of books, cutting the spines to digitize them, and spending tens of millions of dollars on this process. They even faced a copyright lawsuit and settled for $1.5 billion in damages. This shows how valuable high-quality data is.

5. ByteDance’s 10-Trillion-Parameter Dream: Data as the Biggest Hurdle

Developing a 10-trillion-parameter model without relying on pre-trained models poses significant challenges:

  • Exponential Data Needs: The amount of training data required for such a model would be several times that of current models; where will it come from?
  • High Data Quality Standards: Super-large models need precise, professional, and unique content, which requires substantial manpower and funding for screening and annotation.
  • Copyright Risks: If ByteDance follows Anthropic’s approach of scanning physical books, it may encounter copyright issues. This is a problem that the new “Security” department will need to address.

In summary, this organizational change by ByteDance reflects the intense competition for data in the large-model industry. As computing power and algorithms are no longer the only barriers, companies that control high-quality data will have a competitive advantage in future AI developments. ByteDance is taking early steps to prepare for this battle.