Summary of Key Points
Physical AI is becoming the next revolutionary wave in intelligent driving, transforming AI from "thinking" in the digital world (such as generating text and videos) to "acting" in the physical world (such as understanding causal relationships and making proactive decisions). 2026 is dubbed the "Year of Physical AI," as the intelligent driving industry evolves from early modular systems and stress-response-based driving towards end-to-end integration that can predict the future, surpassing human intelligence. The technological approach is shifting from a single VLA (Vision-Language-Action) or world model to a deep integration of both. The competitive focus has also shifted from simply accumulating components (like the number of lidars and computing power) to building a solid foundation in understanding and acting in the physical world. However, we are still far from the ultimate form of physical AI, facing challenges such as insufficient data and high computational demands.
I. Physical AI: The Leap from “Seeing the World” to “Understanding the World”
What makes physical AI different from digital AI? Simply put, digital AI operates in a virtual environment—examples like ChatGPT generating text or Midjourney creating images occur within a digital space. In contrast, physical AI enables machines to understand the movement of objects and causal relationships in the real world and take proactive actions.
For instance, traditional intelligent driving systems might stop at a red light (reacting to what they see), but physical AI could predict that a vehicle ahead might suddenly change lanes, prompting it to slow down in advance (anticipating future events and acting accordingly).
Why is physical AI such a big deal? Elon Musk, the CEO of NVIDIA, calls it the "trillion-dollar growth wave" following digital AI. Success stories like the mass production of the Momenta R7 and the shipment of Yuzhu robots demonstrate that AI is transitioning from the virtual to the real world. For intelligent driving, physical AI solves a critical issue: previous systems lacked the common sense that humans have, but now they can develop a kind of "physical intuition."
II. Intelligent Driving Technology Paths: From “Insect Intelligence” to “End-to-End Integration”
The evolution of intelligent driving is about making machines increasingly human-like:
1. Early “Insect Intelligence” Phase: Relyed on rules and high-precision maps, such as slowing down at speed limits or following predetermined routes. This approach was akin to insects—reactive but unable to handle unexpected situations (e.g., a cat suddenly appearing on the road). Perception, decision-making, and control were separate processes, leading to errors (e.g., failing to respond promptly to obstacles).
2. End-to-End Revolution: The Tesla FSD V12 in 2024 marked a turning point, integrating perception, decision-making, and control into a single neural network, eliminating 300,000 lines of manual code. In simple terms, it processes information directly (e.g., “seeing an image → issuing commands to the steering wheel/accelerator”) without intermediate steps. Domestic manufacturers are following this trend, with phased end-to-end implementations in 2024-2025 and full-end-to-end integration becoming mainstream by 2026.
3. **From “Black Box” to “Explainable”: Early end-to-end systems were “black boxes”—it was unclear why the machine stopped, making it difficult to determine responsibility in accidents. VLA (Vision-Language-Action) was introduced to convert visual information into text (“a pedestrian is crossing the road”) for better decision-making, but this approach overemphasized text understanding at the expense of driving skills (like a student memorizing answers without applying knowledge).
It later became clear that VLA alone was insufficient; a “world model” was needed to simulate real-world behaviors (e.g., predicting that a vehicle ahead would turn left or that braking distances increase in wet conditions). The industry consensus now is that VLA handles the present, while the world model predicts the future. The combination of both is the ultimate solution.
III. Major Industry Shift: VLA + World Model Becomes the Standard
There was once a debate between proponents of VLA and world models:
- Momenta and Huawei initially favored world models, with Cao Xudong suggesting that VLA wasted resources, while Jin Yuzhi argued that world models provide a better understanding of the physical world.
- Xpeng also used VLA but shifted to integration in 2026, using its X-Mind framework to anticipate future events before taking action.
The Tesla FSD V14 further clarified this debate by integrating the Grok large model (related to VLA) with a world model for training. Almost all leading manufacturers are now adopting this approach: NIO has introduced a world model architecture, Horizon’s HSD V2.0 uses a “world model + end-to-end reinforcement learning,” and Li Auto merged its foundational model teams to support intelligent driving, infotainment, and robotics.
In summary, it’s no longer about choosing one or the other; rather, they all need to be combined.
IV. Changing Competitive Logic: From “Component Accumulation” to “Building a Solid Foundation”
In the past, the automotive industry competed based on factors like the number of lidars (e.g., “3 lidars”), chip performance (“1000 TOPS computing power”), and range. In the era of physical AI, the focus is on building the best foundation for understanding the physical world:
- Why? Because this foundation is versatile—Tesla uses the same model for both FSD and Optimus (the humanoid robot); Li Auto integrates it for intelligent driving, infotainment, and robotics; Qualcomm combines automotive and robotics under one division. Whoever masters this foundation controls the value chain.
Leading manufacturers are investing heavily in this: Huawei plans to invest over 18 billion yuan in AI research by 2026 and another 70 billion yuan in computing power over the next five years; Tesla has 230,000 H100 chips with equivalent performance; Xpeng’s Robotaxi has 3000 TOPS of computing power. The competition for computational power is intensifying.
V. How Far Are We from Ultimate Intelligent Driving?
Although physical AI is a promising trend, we still have significant challenges to overcome:
1. Insufficient Data: World models require massive amounts of real-world data for training. Tesla’s FSD fleet has traveled 10 billion miles, while other companies have only covered a fraction of this distance.
2. High Computing Demands: Physical AI models consume more power than large language models, and leading manufacturers are investing heavily in computing resources, but smaller players may struggle to keep up.
3. Moravec’s Paradox: AI excels at complex tasks (e.g., playing chess) but struggles with basic physical perception (e.g., determining if a cup will fall). Current intelligent driving systems lack this level of “physical intuition,” which is why many are still cautious about fully autonomous driving.
2026 is the Year of Physical AI, but we are still far from true autonomous driving. However, this revolution is underway, and we are one step closer to a world where we don’t need to drive ourselves.