虎嗅

The underlying logic of AI's self-improvement: an eternal motion machine, or a turning point for civilization?

原文:AI 自我改进的底层逻辑:永动机,还是文明拐点?

Summary of Key Points

The AI industry is betting on what could be described as “the biggest scientific gamble in human history”: Recurrent Self-Improvement (RSI). This concept involves AI being capable of designing and optimizing its own next-generation models, thereby transforming the computational resources invested into exponential growth in capability. However, this gamble is fraught with uncertainties:

  • The RSI concept has not been clearly defined for 60 years.
  • There are logical deadlocks in the validation mechanisms.
  • Closed-loop self-improvement can lead to model collapse.
  • The current industry investment far exceeds revenue (for example, Google’s parent company had negative cash flow).
  • We are currently in a “bubble” phase where the industry’s activities seem unsustainable.

If RSI becomes a reality, humanity’s role could transform from that of creators to mere elements of the environment or even part of history.

I. The Gamble: Spending More Than We Earn, Yet Unable to Stop

The AI industry is engaged in a “gamble that seems unstoppable”:

  • Exorbitant Investment: The total capital expenditure in the industry exceeds that of the Apollo moon landing and the Manhattan Project. In 2026, Google’s parent company’s capital expenditure could reach up to $205 billion, with its quarterly free cash flow turning negative for the first time (by $5.9 billion).
  • Revenue Falls Short: The income generated by AI does not cover the investment costs, and there is no sign of this gap narrowing in the short term. It’s like opening a business: you spend $1 million on renovation and inventory, but you haven’t started making money yet, and the expenses keep rising.
  • Why Can’t We Stop? All players are increasing their bets because giving up means admitting defeat. Whoever quits first will fall behind.

II. What Exactly is RSI? A Concept That Hasn’t Been Clearly Defined for 60 Years

RSI (Recurrent Self-Improvement) is the “core narrative” of the AI industry, but no one has been able to define it precisely in 60 years. It can be categorized into four levels based on the degree of human involvement:

  • L1: AI Generates Training Data (already achieved): For example, using AI-generated text to train subsequent AI models, which is a common practice today.
  • L2: AI Optimizes Training Processes (under research): AI assists humans in adjusting parameters and modifying code, acting like an “intelligent assistant.”
  • L3: AI Designs New Model Architectures (not yet achieved): AI decides on the number of parameters and structure of the model on its own, essentially “drawing its own blueprint.”
  • L4: Fully Autonomous Iteration (never achieved): AI designs, verifies, and improves its next-generation models without human intervention.

We are currently at most L2; there is still a long way to go before reaching L4. For instance, the recently popular Frontis-MA1 model can only help improve existing code but cannot design new models on its own.

III. The Challenges of RSI: Unverifiable Improvements and Closed-Cycle Vulnerabilities

RSI faces two critical issues:

1. Validation Deadlocks:

  • AI’s claims of self-improvement are not credible (self-evaluation lacks credibility).
  • Human approval is necessary for each step of the process, which undermines the concept of true “self-improvement.”
  • Using a more powerful AI to verify the current model? But who will verify that more powerful AI? It’s an endless cycle, similar to the question of “which came first: the chicken or the egg.”

2. Closed-Cycle Self-Improvement Leads to Collapse: Training subsequent models with data generated by AI is like making copies that become increasingly blurred. Experiments at Oxford University have shown that after several rounds of recursive training, models produce meaningless output (such as random words instead of meaningful text). This is a fundamental principle of information theory: the processing of information does not create new content but only consumes it.

IV. The Critical Factor for Open-Rule RSI: Is the Environment Open Enough?

The only viable path forward is “open-loop self-improvement,” where AI obtains new information from the external environment (the internet, software interactions, and the real world). This depends on the assumption that the environment is sufficiently open to support exponential improvement:

  • Limits of a Limited Environment: For example, AlphaGo became a master of Go, but it could only play within the rules of the game; it couldn’t learn to write poetry. The boundaries of the environment determine the potential of AI.
  • Current Environmental Constraints: The environments accessible to AI (code, web pages, text) are still limited, and it’s uncertain whether they can support exponential improvement. It’s like trying to grow taller while only being fed rice—without meat and vegetables.

V. The Future of Humanity: From Creators to Elements of the Environment, or Even Part of History

If RSI becomes a reality:

  • Short Term: Humans will transition from creators to elements of the AI environment. AI will rely on the internet, chips, and software created by humans, making us one of its information sources, much like a tree in an ecosystem.
  • Long Term: If AI can produce its own chips, generate electricity, and extract materials (achieving embodied intelligence), humanity’s role will become that of a historical condition—similar to how life emerged from primordial soup, but no longer essential for AI’s development.

This is not about “human replacement” so much as a shift in our role. We will go from being the main characters in the story to becoming its backdrop. However, since AI still depends on human infrastructure, this is more of a “changing seats” rather than a complete exit from the narrative.

Final Question: Should We Take This Gamble?

No one can be sure. What is clear, however, is that the risks are not just technological failure but also a transformation of our own role in society. We have already made bets without fully understanding the consequences. The criterion for rationality is not the probability of winning, but whether we have a way out in case of failure. And where exactly is that exit?

Perhaps more important than RSI itself is whether we are willing to accept the idea of transitioning from “creators” to mere “backdrops” in the narrative of AI’s development. That would truly be the real gamble at hand.