虎嗅

Chasing General Artificial Intelligence: Multiple World Models Moving Forward Concurrently

原文:追逐通用人工智能,世界模型多头并进

Summary of Key Points

This article focuses on the “world models” in the field of AI, highlighting that current large language models (such as ChatGPT) have significant shortcomings in tasks involving physical reasoning and spatial planning due to their lack of real-world experience. World models, which can simulate the consequences of actions in reality, are considered a key path toward achieving general artificial intelligence (AGI). The article also discusses the cognitive science behind world models, the commercial progress (including capital investment and strategic moves by leading companies), the disagreements in technical approaches (between pixel-level reconstruction and abstract representation), as well as the gaps that still exist between these models and AGI (such as issues with time abstraction and metacognition).

Detailed Explanation

1. The “Fatal Shortcomings” of Large Language Models: Lacking Real-World “Common Sense”

This is easy to understand with a simple example: If you ask an AI what happens when a cup falls to the ground, it might mention terms like “gravitational potential energy” and “stress waves,” but a five- or six-year-old child would simply say “the cup breaks.” Why? Because the core of large language models (LLMs) is to predict the next word based on vast amounts of text data, yet they have never actually seen a cup falling—without real-world experience, they cannot truly understand physical laws. For instance, while LLMs can recite routes for navigation, they would be confused if asked to take a different path (they lack the mental map that humans use). Even everyday tasks like folding clothes are beyond their capabilities; they can only describe them but not understand the reality behind them. This illustrates the fundamental flaw of LLMs: they are experts in handling text, not in understanding the real world.

2. World Models: Systems That Allow AI to “Imagine” Consequences Like Humans Do

The concept of world models originated from psychologist Alan Turing’s work in 1943, which suggested that the human brain contains a “small model” that simulates the consequences of different actions (for example, knowing that a cup will break if dropped). Later theories, such as the “predictive processing theory,” further emphasized that human perception involves continuously predicting the world and then using sensory information to refine our understanding. AI researchers aim to give AI this ability, allowing it to simulate outcomes before taking action, thus avoiding mistakes in reality. Turing famously stated, “An AI without predictive capabilities is not truly intelligent; intelligence lies in using ideas to test possibilities for us.”

3. Capital Rush: World Models Become a New Frontier in AI

In the past two years, world models have become a hot target for investors and tech giants:

  • Li Feifei (a leading AI expert) founded World Labs, which raised $230 million to develop “spatial intelligence” technologies.
  • Yang Likun left Meta to establish the AMI lab, securing $1 billion to focus on world model research.
  • Google’s DeepMind’s Dreamer4 system, after watching many videos from Minecraft, can independently perform tasks like mining diamonds (a process involving multiple steps, such as gathering resources and creating tools).

However, these are still prototypes: World Labs’ Marble is merely a 3D scene generator, and Genie3 functions more like a game simulator—far from being capable of performing household chores.

4. Technical Disputes: Pixel-Level Reconstruction vs. Abstract Representation

There are two main approaches to developing world models, with intense debates:

  • Pixel-Level Reconstruction: Systems like Dreamer4 attempt to recreate every pixel to achieve a highly accurate simulation of the world after an action. The advantage is detailed accuracy, but it requires massive amounts of data.
  • Abstract Representation: Approaches like Yang Likun’s JEPA focus on capturing only the information necessary for reasoning (e.g., knowing that a cup will break without specifying the position of each fragment). This approach uses less data, but the question remains whether abstract representations are sufficient for complex reasoning. For example, V-JEPA2 can learn object movement patterns, but it hasn’t been proven to support intelligent behavior in open environments.

Both approaches have their merits; Harvna (the lead developer of Dreamer4) believes that pixel-level reconstruction can also lead to abstract understanding, while Yang Likun argues that abstraction is more efficient.

5. There’s Still a Long Way to Go Before AGI

Even if world models become highly advanced, there are many hurdles before achieving true AGI:

  • Time Abstraction: AI can only predict short-term outcomes; it cannot consider long-term consequences (e.g., planning for a better career through current efforts).
  • Metacognition: AI lacks awareness of its own knowledge and understanding; it may give incorrect answers without realizing it.
  • Complex Causality and Understanding Human Intentions: Humans can understand complex causal relationships and predict others’ thoughts, but AI cannot yet do the same.

Experts agree that world models are a necessary step toward AGI, but they are not a sufficient condition; achieving true AGI requires a combination of various capabilities.

In Conclusion

World models represent an important milestone in the journey towards AGI, but they are not a panacea. There is still a long way to go before AI can think in a way that closely resembles human intelligence.