虎嗅

Six Leading Experts in Embodied Intelligence Gather Together: The ChatGPT Moment for Robots – How Many Years Left?

原文:六位具身智能大佬同台:机器人的ChatGPT时刻,还有几年?

Summary of Key Points

At the "Intelligent Empowerment and Embodiment Forum" held by Shanghai WAIC, industry leaders in robotics unanimously predicted that the "robotic ChatGPT moment" would occur within 2 to 5 years. However, they have not yet reached a consensus on the specific criteria for this moment. The discussion focused on three key issues: how to establish a data momentum for robots (which requires real-world experience), whether VLA (Visual Language Action Model) has become obsolete (integration is the way forward), and whether full-stack capabilities are necessary (while having comprehensive abilities, business models can vary). This indicates that although the industry has not yet reached a consensus, it has clearly defined the direction of its exploration.

1. The "Robotic ChatGPT Moment": A Time Frame of 2-5 Years, but No Consensus on Criteria

The guests at the forum predicted that this moment would happen within 2 to 5 years, but there was no agreement on what exactly constitutes the "robotic ChatGPT moment." Yao Maoqing from Zhiyuan provided a clear standard: robots should be ready to use out of the box, capable of understanding natural language commands from ordinary users, and achieve a success rate of 70%-80% in common tasks such as picking up objects or organizing items. Other guests, including Zhang Zhengyou from Tencent and Ma Yecheng from Dyna, did not define specific criteria. This suggests that although there is a general agreement on the time frame, there are differences in understanding what constitutes meeting the requirements—some may refer to the ability to complete simple tasks, while others may mean the ability to handle complex scenarios.

2. Data Momentum: Not Collected, but Acquired Through Real-World Use

Robots need physical experience data (e.g., how much force is required to pick up a cup or how to adjust movement when encountering steps), which is different from large language models that rely on existing text data from the internet. The guests discussed several sources of data:

  • UMI (User-Machine Interaction): Low-cost collection of human operation data, where people perform tasks with devices and the system records the actions and visual information for the robot to learn from.
  • Deployment Data: Data generated from robots making mistakes, pausing, and recovering in real-world scenarios (e.g., folding towels in a hotel) is considered the most valuable type of data.
  • Data Pyramid: The base layer consists of free internet videos and first-person perspective data; the middle layer includes data from actual robots performing multiple tasks; the top layer contains deployment data.

The critical question is whether the chicken comes before the egg: without a good model, customers are hesitant to deploy robots; without deployment, valuable data cannot be obtained. Therefore, the data momentum for robots is built through real-world use, not through data collection.

3. "VLA Is Obsolete": A Misconception; Integration Is the Future

VLA (Visual Language Action Model) processes visual information and commands to generate actions directly, whereas WAM (World Action Model) predicts the outcomes of actions before making decisions. The notion that VLA is obsolete has been popular in the industry, but the forum discussions debunked this idea:

  • Ren Zhiyi from PI stated that their VLA model π0.7 is still leading and can be integrated with world modeling capabilities.
  • Dyna shifted to WAM for certain tasks (e.g., those requiring high robustness) but did not deny the value of VLA.
  • Zhang Zhengyou from Tencent pointed out that robots need a layered system: the upper layer handles reasoning (e.g., "cleaning the room today"), the middle layer handles perception (e.g., "picking up a broom"), and the lower layer controls the motors (e.g., adjusting the broom's force). It is impossible to solve all problems with a single model.

Conclusion: VLA is not dead; what has died is the illusion that one model can cover everything. The future trend is towards systems that integrate VLA and WAM.

4. Full-Stack Abilities Are Needed, but Not Always Required for Business Models

Robot companies follow two main approaches:

  • Model-Oriented Approach: Focusing solely on developing models (e.g., PI, Tencent), but understanding hardware is essential during development (Tencent mentioned, "We don't sell the physical robots, but we need to develop hardware to understand the limitations of models in real devices").
  • Full-Stack Approach: Developing both models and physical robots (e.g., Zhiyuan, Sunday) because real-world challenges like reliability, cost, and battery life cannot be addressed by models alone (e.g., joint control and battery capacity limit intelligence).

The future trend is that full-stack capabilities (understanding both software and hardware) are necessary, but business models can vary. There may be companies that are as comprehensive as Apple or platforms that provide models like Android. However, since hardware has not yet been standardized, even if a company focuses on models, it cannot bypass hardware limitations.

5. Time Is Not the Most Important Factor; Self-Acceleration Is Key

The time frame of 2 to 5 years proposed at the forum is for reference only. The true sign of the "robotic ChatGPT moment" will be when data, models, physical robots, and real-world deployments form a self-accelerating cycle: more robots deployed → more data collected → faster capability improvement → lower costs for entering new scenarios. Until then, all timelines are predictions, and all approaches are being tested. What is certain is that the industry is rapidly moving towards practical applications.

In summary, while there is no consensus on the exact timing of the "robotic ChatGPT moment," the industry is clearly moving in a direction where robots become more useful through continuous improvement and integration of various technologies.