Summary of Key Points
This news article highlights recent significant developments in the field of embodied intelligent robots: The method of data acquisition has shifted from simulation and real-device collection to the use of “human data” (data collected through devices worn by humans), leading to a surge in the volume of data and the corresponding processing demands. This has directly increased the demand for cloud computing power. However, there are still challenges such as a lack of real-device data and scarcity of data for specific scenarios. The industry will need 2-3 years to lay the necessary foundation before achieving large-scale adoption, including developing robust models and gathering sufficient data.
Detailed Analysis
1. Why Does Embodied Intelligence Suddenly Require More Computing Power?
The demand for cloud computing power in the embodied intelligent robot industry has risen rapidly this year—Tencent Cloud reports that the related business’s computational needs have increased by 4-5 times, primarily due to a surge in data processing requirements. Data cleaning and collection account for a large portion of this new demand.
Why has the need for data processing suddenly increased? It’s because the industry is adopting new methods for acquiring training data, which has resulted in a significant increase in the amount of data. Cloud providers have also found solutions to this issue. For example, Tencent Cloud uses “tidal computing power” to utilize idle computational resources at night to perform data annotation and processing, thereby reducing the costs for robot manufacturers.
2. The Source of Data Has Changed: “Human Data” Becomes the New Mainstay
Previously, robot training data came from two sources: simulated synthetic data (generated by computers) and real-device data (collected through robot actions). Now, a third approach has emerged: “human data.” Humans wear lightweight devices (such as hats with integrated phones, smart glasses, or VR equipment) to record their movements and the changes in their surroundings, which are then used to train robots.
What makes human data advantageous? It is easier to collect and more efficient to process than real-device data. For instance, while real-device data may only provide effective data for a few hours out of an 8-hour collection period, human data can yield over 6 hours of useful information. The training approach has also changed: previously, 80% of the data came from simulation and 20% from real devices; now, 80% of the data is used for “pre-training” (teaching basic robot actions), and only 20% for fine-tuning. The amount of data has increased dramatically, from a few hundred hours to several hundred thousand hours, with some predicting that it could reach up to 10 million hours in the future.
3. Robots’ Limited Computing Power: Relying on the Cloud
The computing power inherent in robots (referred to as “edge-side” processing) is insufficient for handling larger models. For example, the CTO of Zeroth stated that to enable robots to perform actions with high success rates, model parameters need to exceed 20 billion (20B), which is beyond the capacity of edge-side processors. For long-term tasks (such as chatting with the elderly or delivering water), additional processing capabilities are required, necessitating the use of cloud computing.
Using the cloud also reduces costs; deploying large models on edge devices is expensive due to high hardware requirements, while the cloud allows for shared computing resources. Many manufacturers rely on Tencent Cloud’s storage and computational services, as well as platforms like ClawPro, to develop robot capabilities (such as fall detection).
4. More Data, but Significant Challenges Remain
Although human data has increased the overall volume of data, there are still two major issues that need to be addressed for robots to become truly useful:
- Scarcity of Real-Device Data: Real-device data is essential for accurate robot operations (e.g., picking up bread), but it is difficult to collect. Even a dataset of 1000 hours of real-device data is considered rare.
- Lack of Data for Specific Scenarios: While there is enough data for simple tasks, there is a shortage of data for complex scenarios (such as massage, retrieving items from under cabinets, or refueling vehicles). The maturity of hardware (e.g., dexterous hands) also affects data collection; only recently have there been sufficient data for more sophisticated operations.
5. How Far Is It to True Adoption? At Least 2-3 Years of Preparation Required
Industry experts believe that embodied intelligent robots are still some way from being widely used in factories and homes:
- Interdependence of Five Elements: Data, computing power, models, scenarios, and robot hardware must work together seamlessly. For example, data is related to both the hardware and the functions of the robots, and any missing element can hinder their effectiveness.
- Long-Tail Effect: Similar to autonomous driving, most progress has been made in the initial stages, but the last 20% (e.g., handling extreme scenarios) will take more time. Currently, robot capabilities still fall short of expectations.
- The Data Loop Has Not Yet Started: The industry is still in the process of accumulating foundational data and models. It is estimated that it will take 2-3 years to establish a cycle where more data leads to better models, which in turn generates even more data. The focus should be on developing solid base models first to enable robots to be trained in specific scenarios.
In summary, the embodied intelligent robot industry is making rapid progress, but it has not yet reached a point of explosive growth. This is a critical phase for building the necessary foundation, and only when data, models, and hardware are fully matured will these robots become an integral part of our lives.