虎嗅

Helping robots with data processing: There's plenty of money involved, but the potential for bubbles is also significant.

原文:帮机器人“搞数据”:钱很多、泡沫很大

Summary of Key Points

The biggest bottleneck in the field of embodied intelligence (such as mobile robots) is the lack of sufficient operational data. Robots perform well in tasks like dancing at the Spring Festival Gala or folding clothes in laboratories, but they fail in real-life home and industrial settings due to a lack of comprehensive data on what they see, what actions to take, how much force to use, and how their joints should move. This has led to the emergence of a separate data industry, with massive capital inflows (for example, Guanglun Intelligence becoming a unicorn and Jianzhi Robotics receiving investment from Ant Group). However, this sector faces numerous challenges: diverse but flawed data collection methods, immature business models, a lack of unified industry standards, and inflated valuations that indicate the presence of a bubble.

Why Do Robots Fail in Real-World Scenarios? Lack of “Closed-Loop Action Data”

For robots to learn tasks like opening bottle caps or pulling out drawers, they need complete data that includes visual information, actions, the amount of force applied, and the angles of their joints (e.g., “seeing the bottle cap → reaching for it → applying 3N of force → rotating the wrist by 30 degrees”). Such data is virtually non-existent on the internet and must be generated specifically for robots.

In the past two years, the industry has focused on developing hardware (robots and dexterous hands). Only after the hardware was created did it become clear that changing the type of cup or surface would cause the robot to stop functioning. Robot manufacturers found that their own data collection efforts were insufficient (for instance, Tesla’s Optimus robots cost $25–48 per hour to collect data), leading to a surge in demand for third-party data services.

How Is Data Generated? Four Methods with Their Own Advantages and Disadvantages

There are mainly four approaches to generating data, and leading companies often use multiple methods:

1. Real-Robot Data Collection: Humans wear VR gear to control robots and perform actions, providing high-quality data that can be used directly. However, this is expensive and can cause dizziness in the operators (with a failure rate of around 30%). Companies like Zhiyuan and iRobot use this method, but some domestic firms have simplified it by making the human arms more similar to robot arms, sacrificing certain specialized capabilities.

2. Data Collection Without Robots: Human actions are used as a proxy for robot movements, reducing costs by half. For example, companies like Mifeng Technology use custom grippers, or Jianzhi Robotics uses motion capture devices. The downside is that it’s difficult to accurately transmit the amount of force needed during tasks like opening bottle caps.

3. Simulation and Synthesis: Data is generated in a virtual environment (e.g., Guanglun Intelligence’s self-developed simulation system, which has produced over 1 million hours of data). This method is cost-effective and scalable, but there are discrepancies between the virtual and real worlds (e.g., virtual cups have fixed weights, while real ones may contain liquid), which can lead to operational errors.

4. Video Analysis: Actions are extracted from YouTube cooking videos and analyzed to understand how hands move. This method is very cost-effective and generates large amounts of data (e.g., Jijia Vision raised 3.5 billion yuan in funding). However, it only provides visual information and lacks details about force and joint movements, making it suitable as supplementary data.

In summary, real-robot data remains the most valuable for customers. Combining simulation and video analysis methods is the current trend (e.g., using simulation to calibrate real robots).

How Do Data Companies Make Money?

Data must be sold to be useful, but business models are still being developed:

1. One-Time Sale of Data Sets: This is the most common approach, but it’s highly competitive. A data set can cost tens of thousands of yuan, and the quality of the data varies (with a 30% failure rate). Once customers establish their own data collection systems, they may reduce purchases.

2. Sale of Hardware: Companies sell wearable data collection devices or remote control workstations. While this creates a barrier to entry, customers may start collecting data on their own after purchasing the hardware, making it difficult to sustain long-term revenue.

3. Sale of Platforms/Subscription Services: This approach has the highest barriers but is the most challenging to enter, as it requires large-scale data demand from customers (which the market has not yet reached).

A strange phenomenon in the industry is that some companies with annual revenues in the millions are valued at hundreds of millions. Investors focus on “value per customer × number of potential customers × data barriers” when evaluating these companies. For example, Guanglun Intelligence’s orders in the first quarter were 550 million yuan, and other companies’ valuations are based on this ratio.

Who Will Set Data Standards?

There is currently no consensus on data standards, and leading companies are competing to establish them:

  • Guanglun Intelligence is developing an evaluation platform called RoboFinals.
  • Mifeng Technology is promoting a cross-platform data format.
  • Jianzhi Robotics is focusing on industrial-grade data formats.

Setting standards involves conflicts of interest; if one company’s format becomes the standard, other manufacturers will have to pay for its use (similar to paying a toll). Currently, manufacturers are more focused on mass production (e.g., Zhiyuan’s Expedition Robot and Yushu’s G1 robot).

In contrast to autonomous driving, where the environment (roads) has stable standards, embodied intelligence operates in diverse settings (factories, homes, operating rooms), making data even scarcer and harder to monopolize. As a result, independent data companies have a longer window for success, but the pace of commercialization is slower.

Is There a Bubble in the Industry?

There are three signs of a bubble in the embodied intelligence data industry:

1. Emphasis on Data Without Corresponding Investment: Everyone agrees that data is crucial, but customers are reluctant to pay for high-quality real-robot data.

2. Focusing on Low-Quality Simulation Data: Many companies produce low-cost simulation data with poor quality.

3. High Valuations and Low Repurchase Rates: Data companies have inflated valuations, but customer repurchase rates are low, and most rely on funding to survive.

In the coming year, the industry will shift from focusing on fundraising to acquiring customers. Companies without orders will be eliminated, and those that can control data production, retain customers, and establish industry standards will likely become the foundation of the “physical AI era.”

In Summary: The embodied intelligence data sector has a clear demand, but it is still in a phase of rapid growth, with both bubbles and opportunities. Only companies that can solve real-world problems will thrive in the long run.