Summary of Key Points
ByteDance has recently consolidated its AI data teams, which were previously spread across various business units, to establish the AI Data and Security Department (parallel to large model departments such as Seed and Flow). This department is responsible for providing cross-modal data services to all of ByteDance’s large models. The previous head, Fu Yue, has left, and the new leader, Wang Yinglei, comes from the TikTok business unit (where he managed live streaming and platform governance), not from a modeling research background. This change reflects ByteDance’s commitment to a self-developed approach for large models (avoiding shortcuts like “distillation”). Data has evolved from being a behind-the-scenes support system to a core company capability. It also aligns with industry trends: AI data production has shifted from merely collecting data from the internet to actively creating it, becoming a new infrastructure for tech companies. However, ByteDance’s decision to appoint a business manager rather than a researcher as the head of the department raises questions about how effectively data production can stay aligned with cutting-edge research.
Why Did ByteDance Promote the AI Data Team to a Senior Level Department?
In simple terms: Data is now the lifeline of large models, and it can no longer be handled in a decentralized manner.
Previously, ByteDance’s AI data capabilities were fragmented across different units, including Global Data, the DMC (Data Management Center), and AIDP under Flow. However, since Zhang Yiming and Liang Rubo made it clear that they would not rely on small models to copy answers from large models but instead focus on self-development, the importance of data has increased significantly. Self-developed models require continuous, high-quality, and diverse data (such as text, video, code, and world knowledge), which must be coordinated across various business units.
For example, if the Seed model needed data, it might have to request it from Global Data, and if the TikTok model needed data, another team would have to be consulted, leading to inefficiencies. With the consolidation, a “central data hub” has been created where all models can get the data they need, addressing issues such as copyright and privacy.
Why Is the New Leader a Business Manager, Not a Model Researcher?
ByteDance chose Wang Yinglei, who previously managed live streaming and platform governance at TikTok, for several reasons:
1. The data department is a “platform-based organization” that requires neutrality and management skills: This new department needs to serve all models (Seed, Douyin, Coding, etc.). If a model researcher were in charge, they might favor their own model’s needs. Wang Yinglei, with his background in business operations, excels at organizational management and cross-departmental coordination, which is essential for managing a team of thousands with a budget of millions of dollars. After all, data production is now a large-scale endeavor, not a small-scale laboratory task.
2. Security and compliance are critical: Wang Yinglei has experience in managing TikTok’s platform responsibilities, including content governance and security compliance. The data department also handles issues related to copyright and privacy (e.g., whether using third-party content constitutes infringement or if user data is leaked), which are his areas of expertise.
However, this raises a question: As data production becomes more research-oriented (e.g., determining what kind of data is needed to improve models), whether having business managers handle it will slow down the feedback process. DeepSeek allows core researchers to directly specify data requirements to shorten the cycle from identifying issues to producing and training models.
How Has AI Data Production Changed?
Previously, large model data was primarily obtained by scraping the internet. But this is no longer sufficient:
- Internet data is becoming scarce: According to Cloudflare, robot traffic has already surpassed human traffic and could exceed it by 1000 times in five years. The internet is like a finite resource; what’s available contains only “answers” without the “thought processes” (e.g., math problems with only solutions).
- Data needs to be actively produced: Models need whatever they lack—programmers for writing code, lawyers for creating legal cases, and even virtual environments for practice. This means that companies must create their own data (e.g., simulating Excel or browsers for AI training).
This is similar to cooking: previously, you relied on the market for ingredients; now, you might need to grow your own vegetables (annotate data) or breed poultry (synthesize data), or even set up a kitchen for chefs to practice (reinforcement learning environments).
Industry trends are even more dramatic. Recruitment platforms like Handshake are hiring doctors and lawyers to produce data, generating annual revenues in the tens of millions of dollars; Google is spending $1.5 billion on virtual work environment companies; Meta is collecting data from employees’ keyboard and mouse activities to train AI. Data has evolved from a raw material to a customized product.
The Industry Trend Behind These Changes: Data as New Infrastructure for AI Companies
ByteDance’s move is not isolated. The entire industry is reorganizing around data:
- Tencent: Although it hasn’t established a senior-level data department, it is poaching ByteDance’s data experts and collecting feedback from over 50 business units to improve model performance.
- DeepSeek: Half of its core researchers are responsible for data annotation because solving AI problems depends on accurate data labeling.
- Meta: It owns 49% of Scale AI, which specializes in reinforcement learning environments.
In short, the competition in AI has shifted from focusing on model algorithms to the ability to produce data quickly and accurately. Companies that can produce the data needed by models more efficiently will have a competitive advantage in the long run. ByteDance’s decision to promote the data department reflects its investment in this new infrastructure.
Is ByteDance’s Approach an Advantage or a Risk?
ByteDance’s approach is characteristic of its style: using organizational skills to scale operations. However, there are potential risks:
- Will the feedback process become slower? Models change rapidly, and what’s useful today may be obsolete tomorrow. If researchers must get approval from business managers for data, could they miss optimal timing?
- Could there be a disconnect between data production and research? Data production now requires model expertise; if managers don’t understand models, will they know which data investments are worthwhile?
Nevertheless, ByteDance has its rationale: self-developed models require sustained, scaled, and compliant data production, and business managers can ensure these foundations are solid. Balancing scale and innovation is the challenge Wang Yinglei faces.
In summary, ByteDance’s adjustment not only demonstrates its commitment to self-developed large models but also responds to new trends in AI data production. It highlights that in the AI era, data is no longer a free resource; it requires significant investment, organizational effort, and expertise to harness effectively. Those who can manage and utilize data well will gain a competitive edge.