虎嗅

Is the new world model released by Li Feifei’s World Labs a real world model?

原文:李飞飞的World Labs新发布的世界模型,是真的世界模型吗?

Summary of Key Points

World Labs’ Atlas is not an ordinary “drawing tool” but a “world simulator” that can generate structured, three-dimensional worlds usable by machines. It utilizes technologies such as depth maps, point clouds, and 3D Gaussian beams to transform ordinary photos into measurable three-dimensional data (e.g., the diameter of a cup or the distance from a table to a wall). It can also combine real photos with AI to train robots in virtual environments, allowing them to function without the need for actual data. The core breakthrough of Atlas is its transformation from a renderer that merely displays images to a simulator that enables machines to perform tasks, thus bridging the gap between AI generation and industrial applications.

I. Atlas: Not for “Drawing,” but for “Recording Data”

The article distinguishes between three levels of AI systems:

  • Renderer: Responsible only for creating visually appealing images. For example, it generates the corresponding view when you move the directional keys, but it doesn’t know the height of a cup or the distance from a table to a wall (these details are useless for machines).
  • Simulator: Keeps track of the world’s state—information like the location of objects, the size of tables, and whether a door is open or closed. This data can be retrieved at any time, allowing robotic arms to pick up objects accurately.
  • Planner: More advanced, it not only knows where objects are but also determines the best time and angle to pick them up.

Atlas falls into the category of a simulator. For instance, if you take a photo of a kitchen, it can calculate the depth of each pixel and use millions of points to reconstruct the kitchen’s shape. These results can be directly imported into game engines, design software, or the control code for robotic arms.

II. Atlas’ Intelligent Features: Solving Big Problems with “Position Memory”

Atlas’s technological innovations lie in multimodal input and spatial memory:

  • Multimodal input: It processes text, images, camera pose (the position and orientation of the camera at the time of shooting, described by six numbers), and depth maps. For example, if you take photos of a sofa from the door and the window, a regular model would consider them two different images, but Atlas, knowing how many steps you’ve taken and how much you’ve turned, would recognize them as the same sofa.
  • Spatial memory: The previous generation, RTFM, had issues: after exploring a virtual house for half an hour, it either stored too many images and ran out of processing power or forgot the layout (e.g., the position of the table changed when you returned to the kitchen). Atlas uses “positions as indices” – to retrieve the kitchen’s layout, it simply looks for photos taken near the kitchen, regardless of when they were taken (similar to how you can immediately recall a meeting from last week without consulting a calendar).

III. Combining “Realism” with “AI Completion”: A New Approach to 3D Scenarios

There were traditionally two approaches to creating 3D scenarios:

  • Reconstruction: Using real photos to recreate the scene, but any missing parts would remain blank (e.g., the Funes World project, which created backups of real buildings without filling in gaps).
  • Generation: AI creates non-existent parts based on common knowledge.

Atlas combines both approaches: it generates less data when there are many photos and more when there are fewer. For example, with just a photo of a garden, it will add the surrounding buildings; with a photo of a small house, it replaces the missing parts with real ones. For robots, whether the patterns on the walls are generated or real doesn’t matter as long as their positions and sizes are accurate, preventing collisions.

IV. Training Robots in Virtual Worlds: The Core Value of the Simulator

The primary use case for Atlas is robot training:

  • After acquiring the robotics company SceniX, World Labs used the Atlas simulator to train robots. They didn’t need any real data and could directly test tasks like packing, winding wires, and moving test tubes. Some tasks could run continuously for an hour without human intervention.
  • The simulator doesn’t have to be a perfect replica of reality. For example, the friction coefficient of walls in the simulator may differ from the real ones, but as long as Test Case A performs better than Test Case B in the simulator, the same is true in reality. Similarly, if the simulator indicates a risky step, the robot will avoid it in reality (just like wind tunnels used by the Wright brothers, which, though different from real sea winds, helped evaluate wing performance).

This approach solves a major challenge in robot training: real-world mistakes are costly (e.g., a robotic arm damaging objects), but simulations allow thousands of trial and error runs to quickly accumulate learning experiences.

V. From ImageNet to Atlas: The Journey of Li Feifei’s Team’s “World Model”

Li Feifei’s team’s ImageNet in 2009 classified internet images to teach machines to recognize objects (e.g., identifying a cup in a photo). Now, Atlas aims to understand how objects are arranged in the world—where a cup is on a table, where it goes after being moved, and whether it’s the same cup when it’s behind the table.

There are two main approaches to this goal:

  • One approach believes that 3D structures will emerge naturally if enough video data is provided (e.g., Meta’s V-JEPA 2).
  • Another approach (Atlas) directly incorporates camera positions, geometric information, and physical constraints into the model, giving the world a structured foundation from the start.

Atlas’s approach is more practical: 3D data is not as abundant as text (there’s a vast amount of text on the internet), and spatial relationships and interactions are difficult to record. Therefore, providing explicit rules to the model is more efficient.

VI. Making the Virtual World Interactive

The article concludes by comparing Atlas to Morrell’s invention: Morrell’s machine could perfectly reproduce a week’s events but couldn’t perform any actions that hadn’t happened before (e.g., knocking over a cup). Atlas, on the other hand, aims to make the virtual world interactive—allowing robots to reach out, touch objects, and make mistakes. This is what makes a true “world simulator.”

For ordinary people, Atlas may not have a direct impact on daily life immediately, but it will make fields like robotics, industrial design, and digital twins more efficient. For instance, robotic arms could be trained in the virtual world without the need for repeated trials in real factories. This marks a crucial step in the evolution of technology from being visually appealing to being practically useful.