第一财经

"Running is just the entry ticket; what's next for robots?"

原文:跑步只是入场券,机器人下一步拼什么?

Summary of Key Points

This news article focuses on the current state of development in humanoid robots and embodied intelligence: the speed of robot manufacturing (one robot produced every 15 minutes) and the speed of action learning (learning to run in 12 hours) have significantly increased, but the level of intelligence is still insufficient. Embodied intelligence is far from reaching the "GPT moment" (a breakthrough in capabilities like ChatGPT), with the main bottlenecks being a lack of data, models, and computing power. Additionally, the industry faces issues with diverse technical approaches and inconsistent standards. There are currently two main development strategies: "accumulating data and computing power" or "understanding the purpose of actions." Scaling up production requires solving challenges related to open source, standards, and commercialization, with different scenarios expected to be ready in 3 to 8 years.

1. Fast Manufacturing and Quick Action Learning, but Still Lacking Intelligence

The "hardware speed" of humanoid robots has become impressive: it took more than 20 years to produce a robot, but now a factory near the exhibition venue can produce one in just 15 minutes, and learning to run takes only 12 hours. However, this is just superficial; robots merely mimic human actions without understanding the purpose behind them. For example, when cutting bread, different people use different techniques, but robots only replicate the actions without knowing the end goal (which is to make a sandwich). This type of "imitative learning" has accelerated industry development but limits scalability, as robots would need to re-learn for each new task, meaning their intelligence remains at the level of performing actions rather than understanding their purpose.

2. Why hasn't Embodied Intelligence Reached the GPT Moment? Three Major Bottlenecks and Data Challenges

People are wondering when embodied intelligence will experience a sudden breakthrough like ChatGPT. The reasons are threefold:

1. Small Models: Current embodied intelligence models have only around 7 billion parameters, compared to the hundreds of billions in early GPT models, a difference of more than ten times.

2. Insufficient Computing Power: Training embodied models requires just a few GPUs, while GPT needs tens of thousands.

3. Lack of Data: Language models can use data from the internet, but robots need data from the physical world (e.g., grasping soft objects with robotic arms or adapting to different lighting conditions). Collecting this data is costly and incomplete (e.g., in extreme weather). Experts believe that embodied intelligence is at least 5 years behind language models due to the lack of sufficient data, large models, and computing power.

3. The Two Development Approaches: Accumulating Data and Computing Power vs. Understanding Intent

Faced with these bottlenecks, the industry is divided into two camps:

  • The "Scale-Up" approach: Continuously increasing the amount of data, models, and computing power. This includes using "synthetic data" (simulated extreme scenarios with GPUs) to complement real data and expanding model parameters, hoping for a sudden improvement in capabilities like with GPT.
  • The "Understand Intent" approach: Focusing on making robots understand the purpose behind actions. For example, the robot developed by Gordon Cheng's team observes humans cutting bread and learns that the purpose is to make sandwiches. This approach reduces the need to learn many actions and allows for more efficient knowledge transfer to other robots.

4. Challenges to Scaling Up: Inconsistent Standards and Redundant Research

Even if technical breakthroughs are made, robots must overcome issues before entering factories and households:

  • Diverse Technical Approaches: Different companies use different hardware and software, leading to repeated development of the same solutions.
  • Data Isolation: Companies often use different data formats, making data sharing difficult and resource-wasting.
  • Commercialization Difficulties: Many robots are still at the prototype stage, with high maintenance costs and a high degree of customization, making mass production challenging.

Therefore, the industry is moving towards open source and standardization, such as establishing humanoid robot open-source communities and sharing operating systems and algorithms to reduce redundant research.

5. When Will Scaling Up Be Possible? Varying Times for Different Scenarios

Experts have different estimates for when these technologies will be ready, but depending on the scenario:

  • Semi-structured Factories: With relatively predictable processes, large-scale deployment could happen in 3 to 5 years, as robots can initially perform tasks well and then be gradually improved.
  • Industry and Home Use: Industrial applications will take 5 to 8 years, while home use will take even longer due to complex ethical, safety, and other considerations.
  • Chips vs. Software: Chips may reach suitable performance and cost in 3 to 5 years, but the software ecosystem will take longer to mature.

In summary, the key to future progress in humanoid robots lies in solving systemic issues such as system coordination and standardization, rather than focusing solely on superficial metrics like running speed and jumping height.

This news highlights that while the hardware speed of humanoid robots has improved rapidly, true breakthroughs will come from advancements in intelligence and industry collaboration. When will we see robots that can cook and clean in our homes? It may still take a few years, but the direction is becoming clearer.