虎嗅

Atlas Paves the Way with 3D Reconstruction: World Models Enter a New Era of Measurability

原文:Atlas用3D重建开路,世界模型进入可测量的新阶段

In-Depth Analysis: Are World Models Just “Dreaming” or “Building Real Paths”? – Taking World Labs’ Atlas as an Example

Hello everyone, I’m your financial journalist and economist. Today, we’re going to discuss a topic that’s causing a huge buzz in both the tech and capital circles: World Models.

If you follow AI, you might have heard of Sora, the model that can generate realistic videos, as well as various news about robots. Recently, a company called World Labs released something called Atlas, which looks a bit like Sora, but it serves a completely different purpose.

The main point of this article is quite insightful: The current race in the world of world models is shifting from “who can create the most visually appealing images” to “who can calculate the most accurately and apply the models practically.” Many companies are still focusing on improving image quality, but what’s really needed for industrial implementation is a digital environment where robots can directly perform tasks.

Below, I’ll break down this complex news into five key aspects in plain language to help you understand the logic and opportunities behind this technological transformation.

---

1. The Core Difference: Sora Creates Videos, Atlas Builds Real Spaces

First, let’s clear up a major misconception: Atlas is not a video generator; it’s a 3D scene reconstructer.

  • Sora (Video Generation): You give it a description, like “a cat running on the grass,” and it produces a video. Its focus is on whether the pixels are connected smoothly and whether the image looks realistic. The final product is a video file for human viewing.
  • Atlas (World Model): You provide it with several photos taken from different angles, and it reconstructs a 3D space. Its focus is on the accuracy of the coordinates and the geometric structure. The final output is a point cloud (a collection of points with coordinates) or 3D Gaussian splashes (ellipsoids with shape and color).

Popular Metaphor:

  • Sora is like creating a stunning oil painting; no matter how realistic it is, you can’t enter the painting.
  • Atlas is like building a real room with Lego blocks. The room has dimensions, and robots can enter it; programs can directly read the data.

Why This Matters:

Robots don’t need to “watch” videos; they need to know where the walls are, how tall the tables are, and whether they can move through them. Atlas provides the “0s and 1s” that robots can process directly, which is the essential foundation for embodied intelligence.

---

2. Technological Breakthrough: From Professional Mapping to Handy Smartphone Shots

In the past, creating a 3D environment for robot training was extremely expensive. You needed professional lidar, dozens of cameras, and hours of manual modeling by engineers. It was like having to hire a surveying team to measure the land before building a house—slow and costly.

Atlas has significantly lowered the barriers and increased efficiency:

  • Minimal Input: You only need at least 2 ordinary smartphone photos, or up to 25, to reconstruct a usable 3D scene.
  • Intelligent Inference: Official demonstrations show that with a close-up of a desk, it can infer the entire room; add a panoramic view, and the position of the chairs is correct; add a side view, and the display is complete. It’s not just simple puzzle-solving; it understands the spatial relationships between objects.
  • Amazing Results: The main courtyard of Stanford University can be depicted from a high-altitude view using just a few ground-level photos taken with a smartphone.

Business Implications:

This means that 3D scene construction has moved from a professional level to a consumer-level one. Ordinary engineers or even non-technical people can use their phones to take a few photos and generate a virtual environment for robot training in minutes. The cost of data acquisition has plummeted, which is a crucial prerequisite for industry explosion.

---

3. Core Capability: Enabling Robots to “See” and “Preview” the Future

Atlas can not only convert photos into 3D models but also control the perspective and timing, setting it apart from ordinary 3D modeling software:

  • Controllable Camera Generation: You can tell Atlas, “I want to view this scene from a 45-degree angle on the left, then the camera should zoom out slowly.” It will generate a video with a resolution of up to 1440p and a duration of up to 1 minute.
  • Advantage: Other models rely on text descriptions for camera movement (e.g., “move the camera to the left”), which can be error-prone. Atlas directly inputs coordinates, ensuring high precision. In blind tests, Atlas outperforms mainstream video models by 75%-94%, especially in complex camera movements.
  • Time-Space Simulation: This is the coolest feature. By recording a video with multiple smartphones, Atlas can “freeze” time, allowing you to review a moment from any angle.
  • Previously: This required dozens of synchronized cameras and cost hundreds of thousands of dollars.
  • Now: Just a few smartphones and Atlas can achieve this.

Value for Robots:

This is more than just for fun. In robot training, we need to know what the cameras will see if the robot takes a step to the left. Atlas can generate the RGB images and depth data that the sensors should perceive in real-time. Moreover, a real scene can be transformed into thousands of variations (changing lighting, moving obstacles, changing routes) without the need to rebuild the environment.

In One Sentence: You only need to visit the physical world once; the digital world can be used countless times. The subject of data collection has shifted from real robots to virtual robots.

---

4. A Critical Shortcoming: Atlas Can Only Create the “Skin,” Not the “Bone”

Although Atlas is impressive, the article honestly points out its inherent limitation: It understands geometry but not physics.

  • What It Is: A high-precision visual renderer that can depict scenes and determine the location and appearance of objects.
  • What It Isn’t: A physics engine. It doesn’t know if an object is soft or hard, how much it will roll when pushed, or whether it will break. These physical properties are invisible in photos and cannot be calculated by Atlas.
  • Technical Barrier: Atlas’s loss function optimizes image fidelity, not physical accuracy. In the simulation process, it only provides visual input; the physical calculations rely on external engines like MuJoCo and PhysX.

Popular Metaphor:

Atlas is like a top-notch art director who can create a lifelike scene but doesn’t understand mechanics. It knows the wall is red, but it doesn’t know if it can support weight. If a robot hits the wall, Atlas can show the visual impact, but it can’t calculate the force or whether the robot will be damaged.

This is Why World Labs Acquired SceniX:

SceniX specializes in adding physical properties (mass, friction, elasticity) to virtual objects. World Labs (geometry) + SceniX (physics) = a complete “Real-to-Sim-to-Real” (from reality to simulation to reality) closed loop.

---

5. Industry Outlook: The Final Mile from “Perception” to “Understanding”

Let’s look at the overall industry trend:

  • Current Stage: Most companies are still competing in geometric reconstruction and visual generation. The winner is the one that can create the most realistic 3D scenes from photos.
  • Future Competition: The real challenge lies in unifying physical properties.
  • Geometry can be inferred from photos using algorithms.
  • However, physical parameters (such as cable damping, fabric texture, joint friction) must be calibrated through real physical interactions.
  • Future world models need to integrate spatial structure and physical rules into a single model.

Implications for Investors:

1. Focus on “Sim2Real” (simulation to reality) capabilities: Don’t just evaluate how impressive the demo videos are; check whether robots can transfer skills from simulation to real robots without any training data and operate stably. World Labs’ demonstrations of robotic arm operations on cables and manipulating deformable objects, running autonomously for an hour without human intervention, show true technical barriers.

2. Data Cost is the Core Barrier: Companies that can generate high-quality training data using the lowest cost (smartphone photos) will lower the R&D barriers for robotics companies, thus gaining a competitive advantage.

3. The Integration of Physics Engines and Visual Models is the Next Trend: Companies that focus solely on vision or physics engines may face integration risks. Companies like World Labs, with both vision and physics capabilities, are more likely to become the “operating system” for embodied intelligence.

In Summary:

The emergence of Atlas marks a crucial step in the transition of world models from academic concepts to practical applications. It solves the problem of “what the world looks like,” but the problem of “how the world moves” still relies on physics engines.

The essence of this competition is not about who can create the most impressive videos but who can provide robots with a more realistic, cost-effective, and usable digital training environment. The day when geometry and physics are truly integrated will mark a turning point for the embodied intelligence industry.