Summary of the Key Points
This was an industry roundtable discussion participated by young scholars from various AI fields at universities in Hong Kong. The entire conversation avoided hype and focused on the practical challenges of implementing the currently popular but poorly defined concept of “world models.” The discussion corrected the common misconception that “AI’s ability to generate videos means it has a world model,” and it also debunked the fantasies that by collecting videos from the internet, a universal robot could be created or that a robot version of ChatGPT would emerge soon. The scholars also pointed out that the often-mentioned “million hours of training” is a misleading statistic. The consensus reached was that the technical approach for world models is not yet established, and it’s not a capital-intensive game solely for large companies; small teams and entrepreneurs still have significant opportunities to make breakthroughs by focusing on specific use cases.
---
Detailed Explanation of the Key Points
1. 99% of the so-called “world models” on the market are ineffective
Many people believe that AI’s ability to generate coherent videos, such as a cat knocking over a water cup, constitutes a world model. However, this is a misconception. Current video generation AI merely arranges pixels to match our visual expectations; it doesn’t understand that the water cup will break or spill when it falls, or that it might be obscured by a table. A truly capable world model should understand the real-world rules. For example, if you ask it to push a cup, it should predict how far the cup will slide and how much water will spill, even if the cup is out of view. Teams working on video generation, autonomous driving, and robotics all use the term “world model,” but they are following completely different approaches and are far from achieving integration.
2. Relying on internet videos to create a universal robot is unrealistic
There was a notion that ChatGPT became a universal model by processing all internet text. Similarly, feeding a robot with videos from platforms like TikTok or Bilibili would create a versatile robot. The scholars dismissed this idea, stating that internet videos mainly show successful, smooth scenarios, such as people holding cups steadily. What robots need to learn are failure cases, which are scarce on the internet. Even first-person videos of household tasks recorded with cameras like GoPro don’t provide useful data for robots, as human hand movements differ significantly from those of robotic arms. Internet videos can only serve as introductory materials; they cannot replace the real-world data needed for robot development.
3. The “embodied intelligence ChatGPT” hype is premature
Many robotics companies claim their robots have achieved “embodied intelligence” after simple tasks like opening doors or pouring water. However, this is not a true world model. ChatGPT’s strength lies in its ability to handle unfamiliar problems, while robots’ current “command-following” capabilities are based on predefined tasks (e.g., opening a door). Robots need to prove their reliability in industrial settings before entering everyday homes, as a failure could result in significant damage.
4. The claimed “million hours of training” is meaningless
Robots companies often boast about millions of hours of training data, but this doesn’t indicate much learning. For example, a vacuum cleaner that repeatedly cleans the same area for 100 hours gains little knowledge. In contrast, language models use unique words as data, but robots can repeat the same actions endlessly without learning anything. The meaningful unit of measurement for robots should be the number of times they perform new or challenging tasks successfully.
5. Future robots won’t require customization for each body
Current robots require significant retraining when their components are changed. A future world model should be able to adapt to different types of bodies and tasks automatically, without the need for extensive retraining. This would significantly reduce development costs and make robots more practical.