Core Summary
Embodied intelligence (robots that can interact physically in the real world) is a hot topic in the tech industry, with massive capital inflows. However, robots still lack versatility—they fail to perform effectively in different tasks or scenarios. The main bottleneck lies in the scarcity of physical interaction data: robots need to learn about the real interactions between humans, objects, and the environment (such as picking up a cup or opening a door), but there is a huge gap in this type of data compared to the vast amount of text available on the internet (by a factor of 1 million to 100 million). Traditional methods of collecting data using real robots are both expensive and time-consuming. The industry is shifting towards using computational power to generate data (via videos and simulations). Although investors remain enthusiastic, they now place more emphasis on practical results and customer repeat purchases, rather than just hearing promising ideas.
Detailed Analysis
1. The Root Cause of Robot Incompetence: Lack of “Real Interaction Data”
For robots to be as versatile as humans, they need to learn from numerous real-world interactions—such as how to handle cups of different shapes or navigate various surfaces. Unfortunately, physical interaction data is scarce. While the internet contains vast amounts of text and images, robots require data from actual physical actions. Currently, leading companies worldwide only have approximately 300,000 hours of high-quality data collected using real robots, but achieving true versatility might require tens of millions or even billions of hours of data. This data gap is 10^6 to 10^8 times larger than that needed for large language models like ChatGPT—it’s like trying to score full marks on an exam with only one question answered.
2. The Challenges of Traditional Data Collection
Traditional methods of collecting data using real robots are inefficient and costly:
- Manual collection: One person can collect a few hundred pieces of data per day, resulting in tens of thousands per month, which is far from meeting the model’s exponential data demands.
- Annotation costs: Training robots for complex tasks (e.g., cooking) requires manually annotating thousands of examples, and the cost increases significantly with more tasks. This approach is akin to trying to plant a forest by planting just one tree per day—how long will it take to achieve a full forest?
3. A New Approach: Generating Data with Computational Power
The industry is adopting a new strategy: using computational power to generate data faster and more efficiently. They are leveraging internet videos (e.g., cooking videos on platforms like TikTok) and custom simulators to create data. For example, RoboScience uses the path of an object moving from one place to another as a standard for data collection, setting up fully automated data production lines:
- Cost: Each piece of data costs just a few cents, a fraction of what it would cost using real robots.
- Capacity: The amount of data that can be generated is virtually unlimited with sufficient computational power. They plan to accumulate tens of millions of hours of video data and terabytes of simulation data this year, which is nearly one-tenth of the amount used by ChatGPT. This approach is like planting trees in bulk, leading to rapid progress.
4. Investor Attitudes: Enthusiastic but More Pragmatic
Although the field is booming, investors are more selective.
- Financing in the first half of the year: There were 288 financings in China, totaling 46 billion yuan, with 49 companies receiving multiple rounds of funding.
- Practical Requirements: Investors are focusing on teams that can address core issues (e.g., the data gap) and demonstrate real-world applications (e.g., whether robots are being used in factories or homes) as well as customer repeat purchases. Simply discussing the concept of “versatile robots” is no longer enough.
In One Sentence
Embodied intelligence is a promising field, but to make robots truly versatile like humans, we need to bridge the significant data gap. The current approach of using computational power to generate data shows promise, but we are still some way from a breakthrough. Investors are becoming more pragmatic. When considering this area, it’s important to look beyond valuations and focus on whether real problems are being solved and practical applications are being realized.
---
Please note that the translation has been adapted to fit the target audience and language style, while maintaining the accuracy and clarity of the original financial news analysis.