虎嗅

29 CEOs of Embodied Intelligence Discuss "Generalization": Consensus and Disagreements

原文:29位具身智能CEO谈“泛化”:共识和分歧

Summary of Key Points

At the 2026 World Robot Conference, 29 CEOs from embodied intelligence companies discussed the concept of "generalization" and reached two consensus points: generalization is the biggest technical bottleneck at present and improving generalization requires data and real-world scenarios. However, there were disagreements on the technical approaches to achieve strong generalization. The market no longer values "imaginary possibilities of generalization" but rather prices products indirectly through four indicators: shipment volume and order quality. Although companies have different interpretations of what generalization means, the causes of the bottleneck, the technical paths, the data needed, and the pace of implementation, they all agree that generalization is crucial for the commercial success of embodied intelligence.

1. What exactly is "generalization"? It's essentially a robot's "learning ability"

In simple terms, generalization refers to a robot's ability to learn from one example and apply that knowledge to new tasks or scenarios without the need for retraining. CEOs provided more concrete explanations:

  • Gao Jiyang from Xinghai Tu: Generalization equals training cost. The shorter the training time, the lower the cost, and the higher the robot's generalization ability. For example, if a robot can learn a new task in 10 hours now, the goal is to reduce that time to 1 hour or even just a few pieces of data by next year.
  • Jia Kui from Kuaiwei ZhiNeng: Generalization can be divided into two parts: semantic and physical. Semantic generalization means the robot knows how to handle different objects (based on real human data), while physical generalization means the robot can handle objects reliably in different environments (using simulation data).
  • Hu Luhui from ZhiCheng AI: Generalization requires the robot to be capable of handling tasks in different contexts (changing tasks, environments, and robots), and all three aspects must be addressed.
  • Federico from Arm: From a system perspective, generalization means the robot can repeatedly use its skills across different objects, such as handling various packages in logistics, which reduces the cost of system modifications.

2. Why is generalization the biggest bottleneck? Robots struggle to adapt to new scenarios

Companies often complain that robots perform perfectly in fixed scenarios but fail in slightly different ones:

  • Wang Xingxing from Yushu Technology: Training robots with data from fixed scenarios results in poor performance in new scenarios due to a mismatch between the robot's understanding and the real world. Language models may not have errors in input and output, but even small deviations in the robot's movements can lead to significant issues.
  • Cao Peng from JD: Logistics robots are efficient in standard tasks but cannot be adapted to complex scenarios (e.g., different warehouse layouts), which hinders scalability.
  • Xiong Rong from Zhejiang Innovation Center: The industry often requires retraining for every new situation, leading to increased data collection and potential loss of precision.
  • Han Zheng from Sudu Technology: Generalization must first achieve 99% success rate; otherwise, the technology remains in the research stage. For example, a robot must be able to hold a cup securely on different tables before it can learn to handle different types of cups.

3. Technical approaches to improve generalization: VLA imitation learning vs. world models that understand physics

Companies fell into two camps on how to improve generalization:

  • VLA approach: Robots learn by imitating human actions (similar to learning skills from videos). However, this approach has limitations:
  • Chen Jianyu from Xingdong JiYuan: VLA can teach robots to handle different cups, but they may not learn new tasks they haven't seen before (e.g., putting a cup in the fridge).
  • Zhang Yufeng from WuJie DongLi: Accumulating data through VLA does not lead to sudden improvements in generalization.
  • World model approach: Robots are taught to understand physical principles (e.g., water flows downhill, and a spilled cup indicates it's upside down). This approach offers better generalization:
  • Chen Jianyu from Xingdong JiYuan: World models can achieve task-level generalization (learning new tasks without examples).
  • Zhu Zheng from Jijia ShiJie: World models should be as versatile as ChatGPT.
  • Song Bin from Beijing FeiDu: Although it will take many years to perfect them, the foundational capabilities must be universal (similar to the nine-year compulsory education).
  • Intermediate approach: Some companies use base models to quickly adapt to new scenarios (e.g., 3 months for overseas applications, compared to 2.5 years for dedicated models).

4. Data is essential for generalization: How much is enough?

All companies agree that data is crucial for generalization, but the amount and type of data required vary:

  • Cao Peng from JD: Training robots for large-scale use requires millions of hours of data; currently, only hundreds of thousands of hours are available. The solution is to collect 10 million hours of human data and 1 million hours of real-world data over two years to achieve better generalization.
  • Xie Chen from GuangLun ZhiNeng: The goal is 1 billion hours of embodied data for high-level generalization, with only 0.1% coming from real robots and the rest from simulations and human data (e.g., the EgoSuite dataset).
  • Wang Xiaogang from SenseTime: World models are used to generate diverse data to complement limited real-world data.
  • Du Dalong from Octopus Power: Hardware advancements (e.g., the EgoBio bracelet) enable zero-data generalization across individuals, allowing immediate use by most users.

5. Implementation pace: Start with practical tasks; the home market is the ultimate goal?

Companies agree that robots with insufficient generalization should not be launched in the home market; they should start with simpler applications:

  • Huang Qingqiu from Moqi Intelligence: There is a trade-off between efficiency and accuracy in industrial applications versus generalization in home use. They plan to start with hotel cleaning tasks, where generalization and data collection are more feasible, before moving onto the home market. Current models can fold clothes, but not handle coats and suits.
  • Huang Yuanhao from Aubi ZhongGuang: They aim to handle dirty and tedious tasks first, then gradually expand to more scenarios, with simultaneous improvements in data, models, and hardware.
  • Qian Dongqi from iRobot: The "bigger the language model, the smarter it gets" rule may not apply to physical AI. They plan to combine human and robotic intelligence to evolve robots from tools to companions.
  • Yu Yinan from Weita Power: Risk management is important; the cost of failure varies in different scenarios (e.g., a plastic cup vs. a glass cup). Products are first tested in a controlled environment before being released.

6. Disagreements remain, but the direction is clear

Despite differences in technical approaches, all companies agree that generalization is key for embodied intelligence to move from the laboratory to practical applications. The market is also using practical metrics (such as shipment volume and orders) to evaluate the effectiveness of generalization. In the next few years, the company that solves the generalization bottleneck will lead the embodied intelligence industry.