第一财经

Embodied intelligence data is surging, but model validation has not yet been fully achieved.

原文:具身智能数据“狂飙”,但还未完全实现模型验证

Summary of the Key Points in the News

This report highlights the most critical turning point in the current humanoid robotics industry: previously, the entire industry was stuck on the problem of a severe shortage of training data. Now, a group of domestic companies have pioneered the use of "data collection factories" for the industrial production of training data for robots, achieving annual output of hundreds of thousands of hours of high-quality data, which has solved the initial issue of insufficient data collection. However, the industry has quickly realized that the data generated in artificially created simulation environments has a clear limit. To create truly functional humanoid robots, a positive cycle must be established where the robots are first trained to a certain level of proficiency, then deployed in real-world scenarios to collect data that can be used to improve their models. The industry is now at a critical juncture, shifting from competing on the scale of data collection to achieving practical, operational effectiveness.

---

Detailed Explanation of the Key Points

1. Why is the entire industry so eager to collect data? The fundamental reason is that the "professional training materials" for robots were previously unavailable.

The large AI models we use today are trained on existing text and video content from the internet, which essentially summarizes thousands of years of human knowledge. However, humanoid robots need to learn how to perform tasks in the real world—how to hold a cup without dropping it, how to avoid objects while cleaning a table, and how to cross a wire on the floor. There is simply not enough precise data labeled for these actions on the internet. Previously, many robots could not even perform basic tasks like opening doors properly because there were too few examples of different door handles, door weights, and obstacles behind them. The current rush to collect data is about building a comprehensive "training library" for robots; otherwise, humanoid robots would remain mere toys that can only move their limbs randomly.

2. The industry has established "robot-specific driving schools" to efficiently generate training data.

Previously, collecting training data for robots was extremely inefficient. Either the robots had to record their movements randomly, resulting in most of the data being useless, or humans had to film the actions, but the details were often inaccurate. Now, leading companies have set up specialized data collection bases, creating a kind of "robot driving school." These bases replicate real-life scenarios such as home kitchens, supermarket shelves, factory assembly lines, and pharmacy shelves, and hire operators to control the robots to perform various tasks. Each action is recorded with precise parameters for joints, force, and path. AI systems then automatically filter out the ineffective data 24/7, and the data is processed in the cloud overnight, ready for use by the training teams the next day. This industrialized approach has enabled top bases to produce hundreds of thousands of hours of high-quality data annually. Some companies have even accumulated millions of hours of useful data, providing robots with the equivalent of years of expert guidance. The problem of insufficient data collection has been largely resolved.

3. Robots trained in such "driving schools" have inherent limitations due to the limitations of simulated data.

Although data collection factories are efficient, the robots produced are like novice drivers who have passed all the basic tests but lack practical experience. They have three major flaws: the simulated environments are too artificial (no unexpected events), the standards for data collection are inconsistent (incorrectly labeled joint positions can lead to incorrect behavior), and the pace of data collection is too slow (real-world work requires faster movements). There is a consensus in the industry that data collected alone cannot produce truly functional robots.

4. The ultimate solution is to deploy robots in real-world scenarios for continuous improvement.

The best approach is to first train the robots using high-quality data from data collection factories to a sufficient level of proficiency, then deploy them in real-world tasks. For example, having them sort parcels at a courier station or serve dishes in a cafe. While working, the robots collect data that can be used to refine their models. Once this cycle is established, robots' capabilities improve rapidly. The report mentions that robots may start with low accuracy at a courier station, but after collecting enough data, their performance improves significantly. This cycle of "working, collecting data, upgrading, and working better" is the key to breakthroughs.

5. The claimed "million hours of data" are just a starting point; the competition is far from over.

Companies are boasting about the amount of data they have collected, but none can claim with certainty that they can create perfect humanoid robots. The industry is still in the process of figuring out the exact amount of data needed to develop effective models. The millions of hours of data collected so far are more like an entry ticket to the competition. The real winner will be the one that can gather a large amount of data from real-world scenarios. In the next one or two years, we will see more robots working in courier stations, cafes, and factories. These seemingly mundane scenarios will be the true test grounds for the humanoid robotics industry.