Summary of Key Points
This article, by drawing an analogy to the cognitive growth of humans (from memorizing individual cases to abstracting general principles), highlights a common misconception in robot training: overemphasizing the size of the data while neglecting the ability to generalize—i.e., the robot's ability to quickly adapt to new tasks with minimal new data. The article discusses two approaches that robot companies can take to improve generalization: either expanding data coverage or enhancing the efficiency of task transfer. It provides examples of different companies' technical choices and outlines criteria for evaluating generalization capabilities (five specific abilities plus a task cost curve), ultimately concluding that the ability to generalize, which leads to a continuous reduction in task costs, is the key to a robot company's long-term competitiveness.
1. Core Cognitive Shift: Generalization Ability is the “Intellectual Work,” While Data Volume is the “Physical Work”
The article begins with the author’s own growth experience as an example: transitioning from writing opinion pieces to emotional essays to technical analyses. He did not have to learn everything from scratch in each new area but instead abstracted common principles (such as “the essence of a problem lies in its contradictions”), which allowed him to quickly master new fields. The same principle applies to robot training:
- Physical Work Approach: The idea is that more data equals better performance, so companies focus on collecting and labeling large amounts of data to train robots to memorize specific cases (for example, how to grasp a cup on the left side of a table).
- Intellectual Work Approach: The goal is to enable robots to derive general rules from a small amount of data (for example, considering the center of gravity and friction when grasping a cup), so they can handle new situations without relearning.
The author concludes that generalization ability is the crucial factor—just as humans rely on deep cognitive understanding to adapt to new problems, robots rely on this ability to handle new tasks with minimal data.
2. Two Approaches for Robot Companies: “Memorizing More Cases” vs. “Deriving Rules”
There are two contrasting research and development strategies in the industry:
1. Expanding Data Coverage (Physical Work Approach):
- Representative companies: Physical Intelligence, Figure AI, etc.
- Approach: Invest resources in data collection (e.g., using teleoperated robots to collect motion data) and data labeling to train robots to recognize more cases. The logic is straightforward: the more cases seen, the higher the likelihood of handling new tasks successfully.
- Problem: This approach is less unique (many companies can do it), and data can never cover all possible scenarios.
2. Enhancing Task Transfer Efficiency (Intellectual Work Approach):
- Representative companies: Field AI, Mujin, etc.
- Approach: Focus on deriving reusable rules from existing data, such as transferring grasping techniques learned from one task to another. The key question is: “Can we complete new tasks with the least amount of new data?”
- Advantage: This approach is more valuable and cost-effective. For example, Mujin can use existing knowledge to handle new box shapes without re-collecting data, and Field AI enables robots to adapt to changing environments through a “world model.”
3. The “Secret Weapon” of Generalization Ability: Reusable Structures
To achieve generalization, robots need reusable structures—knowledge that can be applied across different scenarios and tasks. The article identifies two types of such structures:
- When Environmental Boundaries are Unclear: Use a “World Model”
- Examples: Ice cream shops (with multiple ordering and topping procedures in a flexible environment) or dangerous outdoor areas (with complex terrain).
- Representative companies: Sharpa (ice cream shop robots), Field AI (outdoor robots).
- Approach: First, thoroughly understand one scenario (e.g., spending 100,000 hours on an ice cream shop), then derive a comprehensive world model that combines sensory data (touch, force, etc.) to reduce adaptation costs for new scenarios. Field AI uses a “belief world model” to continuously update the robot’s understanding of its environment and adjust its behavior accordingly.
- When Environmental Boundaries are Clear: Use “Geometric/Physical Rules”
- Examples: Industrial logistics (unloading and stacking tasks in fixed environments).
- Representative companies: Mujin (Japanese industrial robot company).
- Approach: Incorporate 3D sensing, collision constraints, and motion planning directly into the system. This way, the robot can use existing geometric rules to handle new tasks without re-modeling, resulting in almost zero adaptation costs.
4. How to Determine If a Robot Really Has Generalization Ability
The article provides two criteria to avoid being misled by company marketing:
1. Verification of Five Specific Abilities
Generalization is not a vague concept; it needs to be evaluated in five specific scenarios:
- Changing Objects: Can the robot handle new box shapes or parts without retraining? (For example, Mujin can handle different box types without re-registering them.)
- Changing Scenarios: Can the robot continue to function in different lighting conditions or complex terrains? (For example, a four-legged robot in remote areas can adapt to complex terrain during power inspections.)
- Changing Tasks: Can the robot combine previously learned actions for new tasks? (For example, can a robot transfer its grasping skills to pouring coffee?)
- Changing Forms: Can the same set of strategies be used for different robot types? (For example, whether multiple-configuration robots share control mechanisms.)
- Handling Failures: Can the robot automatically recover from errors (e.g., adjusting grip strength when the cup slips.)
2. Task Cost Curve
The most important indicator is whether the time required to complete tasks and the amount of new data needed are continuously decreasing. For example, Google’s Gemini robot only requires 50–100 demonstrations to learn new tasks due to its prior knowledge base.
- Counterexample: Figure AI claims to train logistics tasks in 8 hours but does not disclose the amount of pre-trained data or subsequent task costs, so its generalization ability is questionable.
Only a continuously declining cost curve indicates that the robot is building reusable capabilities, not just a large collection of case data.
5. The Key to Long-Term Competition: Making Task Costs Lower
The article emphasizes that the core competitiveness of robot companies lies in their ability to reduce task costs over time. Just as humans benefit from the compounding effect of experience, robots benefit from the compounding effect of generalization—each completed task makes subsequent tasks faster and cheaper.
In a few years, the difference will be clear: some companies will complete the 11th task in one day after doing 10 tasks, while others will still need a month. This is the essential difference brought about by generalization ability.
The value of this article lies in its recognition of the misconception that more data always equals better performance. It provides a new perspective for robot companies and investors to evaluate their competitiveness, focusing on generalization ability and cost efficiency rather than data volume. For the general audience, it also clarifies the core competitive factors in the robotics industry, avoiding the misconception of the “big data” hype.