Summary of Key Points
This article reveals the core debate in the field of Physical AI by analyzing the technology of ball-playing robots developed by teams such as MIT, DeepMind, Berkeley, Galaxy General, and Sony AI. The issue is not about which of physical models or machine learning should replace the other, but rather about how the two can work together effectively. Physical models provide stable and low-cost “known knowledge” (such as the laws of gravity and the basic motion of balls), while machine learning handles high-dimensional, complex “unknown skills” (such as the coordinated movement of humanoid robots to hit balls). The different teams’ approaches essentially involve reassigning the responsibilities of these two components: from using models to directly control the robots, to using them to create simulation environments for training, to using them to define “possible worlds.” The ultimate goal is to achieve a hybrid model where physics provides the framework, and learning fills in the details.
Detailed Analysis
1. Why are ball-playing robots a good example of technological approaches?
Playing ball may seem simple, but it actually highlights the most critical challenges in robotics:
- Extremely short reaction times: Table tennis balls move at speeds over 20 m/s, with intervals between hits of less than 0.5 seconds. Robots must predict the ball’s trajectory in advance; otherwise, they will react too late.
- Complex dynamic changes: The rotation of the ball (e.g., 1000 rad/s) can alter its flight path and collision outcomes, and factors like air resistance and ground friction cannot be ignored.
- Full-body coordination: When a humanoid robot plays tennis, it must coordinate more than 20 joints (legs, waist, shoulders, etc.) simultaneously while maintaining balance. There are often dozens of possible body postures for the same hitting action.
These challenges clearly expose the contradiction between using models and relying on learning: trying to calculate all details with models is too complex, while relying solely on learning can lead to errors. Therefore, ball-playing robots serve as an excellent example for observing different technological approaches.
2. The evolution of physical models: from control to simulation, to defining possibilities
Physical models have not disappeared; they have just changed their role within the system:
- MIT: Direct model control (2025 tennis ball robotic arm): The robot uses physical formulas to calculate the ball’s future trajectory after each reception and then plans its swinging motion, similar to an engineer guiding each step manually.
- DeepMind: The model is embedded in simulation (2024 competitive tennis system): The robot trains millions of times in a virtual environment built according to physical laws, developing “muscle memory” before moving to the real world (the model creates the virtual training environment).
- Berkeley HITTER: Hierarchical division of labor (2025 humanoid tennis ball robot): The upper layer uses models to predict the ball’s trajectory and the target position of the racket, while the lower layer teaches the robot how to move (using models for ball dynamics and body movements).
- Galaxy General LATENT: The model defines “possible worlds” (2026 tennis robot): During virtual training, the weight of the ball, ground friction, and air resistance are intentionally varied within reasonable ranges, allowing the robot to learn to adapt to various situations (the model no longer aims for precision but provides a margin for errors).
- Sony Ace: Hybrid approach (2026 Nature paper): Machine learning is the core, but physical simulation training, ball rotation estimation, and collision constraints optimization are still used (learning is the main focus, with the model filling in gaps).
3. Why haven’t physical models been eliminated? Because knowledge reuse is too cost-effective
The core value of physical models is to provide “cheap prior knowledge”:
- For example, the formula for gravity (F = mg) requires almost no cost to derive, but it would take a lot of data and computing power for a robot to learn this through trial and error (e.g., by dropping the ball hundreds of thousands of times).
- Calculating the ball’s trajectory using Newton’s laws is much faster than having the model learn it from videos.
These stable and generalizable rules are more economical to use directly for robots than having them learn them from scratch—just like you don’t teach a child to rederive the concept “1 + 1 = 2”; you teach them directly.
4. The truth behind the debate about different approaches: it’s about dividing tasks and calculating the best options
At its core, this is an economic question: Is it cheaper to use equations, data, or computing power for a particular task?
- Equations are preferred for stable, simple, and low-cost-to-express rules (such as gravity and the basic motion of balls).
- Learning is preferred for high-dimensional, complex, and hard-to-enumerate behaviors (such as the coordinated movements of humanoid robots).
- Hybrid models are used for situations that fall between these extremes (e.g., the complex rotation of balls, where models and real data are combined for correction).
For instance, the hierarchical approach of HITTER uses models for the ball’s trajectory (cost-effective) and learns the body movements (cost-effective). The randomization in LATENT defines the range of parameters (to avoid repeated calibration) and allows the robot to learn and adapt to changes (high robustness).
5. The future of general-purpose robots: Robustness is more important than precision?
General-purpose robots need to operate in various environments (homes, factories, outdoors), where it’s impossible to calibrate parameters perfectly for each situation (e.g., ground friction, ball wear). Therefore:
- We no longer pursue “absolute precision”; it’s more practical to let robots adapt to “possible worlds” (like with LATENT) rather than creating a perfect virtual replica.
- The division of tasks will continue to evolve: As data and computing power increase, tasks that currently require manual input (e.g., complex interactions) may be replaced by learning, and parts that currently rely on learning (e.g., safety constraints) may be enhanced by models.
In the end, Physical AI will converge towards a model where physics provides the framework, and learning fills in the details—without abandoning known rules and fully exploiting the flexibility of machine learning.
In one sentence
The technological evolution of ball-playing robots reflects humanity’s ongoing exploration of what knowledge is worth imparting directly to robots and what they should learn on their own. Physical models and machine learning are not competitors but partners, working together to make robots smarter and better adapted to the real world.