第一财经

"Seizing the Global Market: The Next Round of Elimination in Embodied Intelligence"

原文:抢滩世界模型,具身智能的下一场淘汰赛

Summary of Key Points

Companies from various industries around the world (robots, brain-computer interfaces, internet giants, etc.) are competing for the “world model” – a technology that enables AI to truly understand the laws of the physical world, predict outcomes, and even interact to change reality. While overseas players have divergent approaches, Chinese companies are more focused on practical implementation. Brain-computer interface firms have pioneered a unique path from translating thoughts into actions. Although there are disagreements within the industry regarding technical frameworks, there is a consensus that simply generating videos is not enough. As capital rushes in, challenges such as unresolved pathways, difficulties in implementation, and data barriers remain. The ultimate winner will be determined by a combination of data quality, industry understanding, and engineering capabilities.

1. The Global Race for the “World Model”: AI’s Guide to the Physical World

In simple terms, a world model is what allows AI to understand the rules of the real world just like humans do – for example, that a glass will shatter when dropped, a table can be moved when pushed, and a ball will follow a parabolic path when thrown. Traditional AI systems mainly processed text and images; now, these companies aim to make AI capable of “seeing, understanding, predicting, and even altering” the physical world. The participants come from different backgrounds: robotics companies want smarter robots, brain-computer interface firms seek to convert thoughts into actions, and internet giants have a advantage in data. They are all working towards the same goal: making AI capable of performing tasks in the real world, rather than just theorizing.

2. Divergent Approaches at Home and Abroad: Overseas Players Take Different Paths, Chinese Companies Are More Practical

Overseas players are following different strategies:

  • NVIDIA uses large datasets of videos to train basic models (equivalent to showing AI a vast amount of real-world footage to learn patterns);
  • DeepMind focuses on creating “interactive world generation” (e.g., building virtual environments for AI to practice interactions);
  • Li Feifei’s team approaches the issue from the perspective of spatial intelligence (e.g., helping AI understand room layouts);
  • Tesla invests in enhancing robots’ mobility (first making the robots more agile, then teaching them to understand the world).

Chinese companies are more focused on practical applications:

  • Shengshu Technology is transferring video generation technology to a universal world model, integrating “understanding, prediction, generation, and action” into a single system;
  • Daxiao Robotics targets specific scenarios such as homes, retail, and hotels, combining vision, language, touch, and motion tracking to create robots that can perform tasks directly.

3. Disagreements Despite the Noise: Generating Videos Is Not Enough

There are differences within the industry:

  • Shengshu Technology believes that the underlying principles of video generation and robot control are similar, so a universal framework should be used;
  • Daxiao Robotics emphasizes that physical scenarios cannot afford mistakes (e.g., a robot breaking a glass could cause damage), so “physical reasoning” capabilities must be built into the technology (e.g., knowing how to hold a glass without dropping it).

However, there is a clear consensus: generating attractive videos is just a starting point; the real challenge is to predict the next steps in the physical world. A video generation model might be like an artist, capable of creating realistic scenes but unable to understand that moving a table will have consequences. A world model, on the other hand, should be like an engineer – it must not only understand the rules but also guide actions.

4. Capital Frenzy, But Three Major Hurdles Remain

Investment is booming: In the past two years, more than 370 new intelligent robotics companies have been established in China, with many valued at over tens of billions. Yet, practical challenges are significant:

  • Unclear Paths: Should we develop a universal model or specialized models for specific scenarios? Should we focus on video generation or physical prediction? These questions remain unresolved;
  • Difficulties in Implementation: Demonstration videos may be impressive, but real-world applications can be impractical (e.g., a robot that works well in the lab may fail in everyday use);
  • Data Barriers: Companies lacking unique real-world data (e.g., interaction data from home environments) or the ability to collect large amounts of data may be at a disadvantage.

5. The Final Competition: Three Critical Factors Determine Success

Who will emerge as the winner? The key lies in three areas:

  • Data Quality: Possession of authentic, unique real-world data (e.g., brain-computer interface companies’ data on thought-to-action feedback);
  • Industry Understanding: Knowledge of the actual needs of specific industries (e.g., understanding the problems that home robots need to solve);
  • Engineering Capability: The ability to transform technology into usable products (e.g., integrating a world model into robots that function reliably in real-world scenarios).

Furthermore, as AI develops the ability to make autonomous decisions, safety and governance will become important considerations – for example, how to prevent robots from causing harm to humans. These issues must be addressed in advance.

In summary, the world model is a crucial step for AI to transition from the virtual to the real world. While everyone is competing for a share of this technology, the ultimate winner will be the one that can solve real-world problems and possesses the necessary capabilities.