虎嗅

Decoding 1 Million Hours of Embodied Intelligence: Can We Really Train a “Smart” Robot? | AI 100 Closed-Door Meeting

原文:拆透具身智能100万小时:真能喂出一个“聪明”机器人吗?|AI 100闭门会

Summary of Key Points

The field of embodied intelligence (robotics AI that can interact with the physical world) is currently booming, but there is a misconception in the industry regarding the notion that "million hours of data will lead to the GPT moment." The effectiveness of the data—its information density, coverage of challenging cases, and transferability—is more crucial than the sheer volume. The industry is shifting from simply accumulating data to focusing on efficiency, with an emphasis on practical applications in industrial and commercial services (wheel-based robots are often more useful than humanoid ones). For the time being, there is no need to choose between different model approaches (VLA vs. world models). The competitive advantage for data companies lies in their ability to create a closed loop that includes data collection, training, deployment, and learning from failures, rather than just the amount of data they possess. There is no significant gap between China and the United States; China has advantages in supply chains and hardware.

Detailed Analysis

1. Are Million Hours of Data Just a Buzzword? Effective Data Is What Matters

Many believe that embodied intelligence requires millions of hours of data, but Peng Pai from Coolwha Technology provided an counterintuitive example: driving 100 kilometers on a regular road won’t provide much useful information for the model, whereas a single plastic bag floating in a market can reveal differences in how lidar and vision systems handle flexible obstacles—this illustrates the importance of "information gain."

Key Criteria for Effective Data:

  • The source of the data is unreliable; many claims of millions of hours of data are based on the number of tokens from language models, which cannot be verified or disproven.
  • Effective data should meet three criteria: it must be explainable (failures can be attributed to specific factors such as angle, force, or strategy), have a high density of challenging cases (rare but error-prone scenarios like heavy rain or mixed traffic), and be reusable across different robots (not a one-time cost).
  • The point at which data collection becomes unprofitable is when the business value of new data is less than the cost of acquisition; for example, if 99% success rate can be achieved with existing data, further collection is not worthwhile.

2. The Data Pyramid Has Changed: Human Videos as the Foundation, Real-World Cases as the Essence

The "data recipe" for training embodied intelligence is not fixed and has evolved into a dynamic pyramid:

  • Base Layer: First-person human videos (e.g., cooking or cleaning videos online) are large in scale and useful for training multiple robots.
  • Middle Layer: Simulated or synthetic data is used to simulate extreme scenarios (e.g., nuclear power plant failures) to supplement real-world failure cases that are difficult to collect.
  • Top Layer: Real-world challenges encountered during actual operations are the most valuable, as they directly improve model stability.

Core Logic: These three layers should be integrated in a cycle—for example, if a robot encounters a problem with a plastic bag, simulations can be used to recreate the scenario, and then human videos can provide additional context. Guo Junliang from Qianxun Intelligence added that pure human videos lack details of robot interactions, so UMI (User-Machine Interaction) devices (e.g., handheld grippers) are also used in the training process.

3. Model Approaches: No Need to Choose Between VLA and World Models for Now

There are two main approaches: VLA (learning actions directly from videos) and world models (learning physical laws first before making decisions). Experts agree that when there is enough data, the difference between the two methods is minimal. A layered system, like Google Gemini Robotics, is more practical; higher-level models handle tasks (e.g., understanding the task of cleaning), while lower-level VLA systems handle specific actions (e.g., rotating the brush). In commercial scenarios, customers often don’t need a single model to perform all tasks.

4. Commercialization: Start with Practical Applications

The focus has shifted from developing robots that can adapt the environment to designing robots that fit into human societies.

Priority for Implementation:

  • Exclude extreme cases (e.g., highly customized industrial scenarios where traditional automation is already mature and less profitable).
  • Home environments are too complex (with elderly people, children, and pets) and not yet profitable.
  • Practical applications in industries or commercial services (e.g., factory logistics, mall cleaning) offer a balance between customization and generalizability, with minimal human-machine interaction required.

Humanoid Robots Are Not Essential: Guo Junliang from Qianxun Intelligence noted that a wheel-based robot with a lifting mechanism can cover over 95% of tasks; customers care about stability, maintenance costs, and ROI, not whether the robot looks human-like.

5. How Do Data Companies Survive? Quality Data Is More Important Than Quantity

The core competitiveness of data companies lies in their ability to create a closed loop that ensures learning from failures:

  • A closed loop (data collection → filtering → training → deployment → retraining) is essential for continuous improvement. For example, if a robot gets stuck on a plastic bag while cleaning the road, the data company should record this failure so the model can avoid it in the future.
  • Data must be flexible and transferable across different robots; it shouldn’t be tied to a specific model.
  • Data companies need to collaborate closely with model teams to understand their needs and provide relevant data (e.g., repeated data from common scenarios is less useful, while challenging cases are more valuable).

Conclusion

The competition in embodied intelligence has shifted from who collects the most data to who can use it most efficiently. Million hours of data are just the starting point; the real milestone is turning expensive real-world failures into valuable learning experiences for models. In the future, robots may not necessarily look human-like, but they will be smarter and better integrated into our lives—able to avoid obstacles in markets and on sidewalks.