虎嗅

RSI: Will the Strong Really Get Stronger as AI Learns to Evolve on Its Own?

原文:RSI:当AI学会自我进化,强者真的会愈强吗?

Summary of Key Points

In 2026, the most intense technological competition in Silicon Valley will revolve around RSI (Recursive Self-Improvement) – which, simply put, involves using AI to identify problems and optimize itself, creating a cycle of "discovery → improvement → greater strength → re-discovery" that accelerates over time. Giants like OpenAI and Anthropic, as well as the academic community, are all betting on this trend. The central question is whether the giants will dominate through their computational power or if newcomers can make breakthroughs and overtake them. At the same time, security and organizational efficiency have become key factors in this competition. This race not only determines the future of the AI industry but also touches on a deeper issue: whether humans will cede control of their own evolution.

I. What is RSI? Why has it suddenly become so popular in 2026?

The core of RSI is a closed loop where AI continuously improves itself: it identifies weaknesses in its models, optimizes them, becomes stronger, and then uses its enhanced capabilities to find even more profound issues for further optimization. This cycle accelerates rapidly.

Its current popularity is not accidental; two conditions have come together:

1. AI is now capable of taking on tasks independently: For example, it can write code, run experiments, and analyze results on its own.

2. AI understands itself: It can analyze its own weaknesses in a way similar to how researchers would.

More importantly, the feedback cycle has been compressed to minutes. In the past, human researchers would spend days or even weeks from conceiving an idea to obtaining results (through a mentor-student relationship and experimental processes). Now, AI can complete these steps in just minutes, akin to switching from a bicycle to an electric vehicle for faster progress.

Major players are already taking action: OpenAI has set a timeline for AI research internships in 2026, Anthropic’s Claude has independently completed 800 hours of AI security research (four times more efficient than humans), and the ICLR conference dedicated a special session to RSI – indicating that it has moved from being a niche topic to a industry-wide consensus.

II. How does RSI differ from AutoML from 10 years ago?

Some might ask: Didn’t Google’s AutoML also involve AI optimizing AI back then? The difference is significant:

  • In the past, AI was "confined": Humans defined the boundaries (such as search scopes), and AI could only find optimal solutions within those limits, like a bird in a cage.
  • Today’s AI is semi-wild: Large models can understand their own architecture and problems without human guidance, expanding the scope of exploration and finding new directions for optimization.

In other words, while humans used to control the details, AI now assists them in identifying potential paths; the role of humans has shifted to making higher-level decisions.

III. The central debate: Will the giants dominate, or will newcomers have a chance?

This is the most intriguing aspect, with two main viewpoints:

  • The "Power Leads to Victory" camp believes that those with more computational power and coding skills will accelerate the RSI cycle, leaving others far behind (e.g., giants like OpenAI and Anthropic).
  • The "Breakthroughs Are Unpredictable" camp argues that intelligent improvement follows a stepwise pattern of rapid growth, followed by periods of stagnation, with breakthroughs coming from new principles (such as Newton’s discovery of gravity). This means smaller teams could also make breakthroughs if they stumble upon these new principles first.

The disagreement reflects different beliefs about how AI will evolve: one side believes in the "scaling law" (more power equals greater performance), while the other emphasizes that breakthroughs rely on inspiration. One of Anthropic’s co-founders even set 2028 as a benchmark for proving whether complete automation of RSI is possible.

IV. Where does RSI stand now? What is still lacking?

RSI is still in its early stages:

  • Achievements: AI has already shown impressive capabilities in tasks that require scoring and rapid feedback (e.g., algorithm optimization and efficiency improvement). For instance, Claude from Anthropic wrote 80% of the company’s code, and engineers are now eight times more efficient than in 2024. The OpenRSI model from Tsinghua University has approached the performance of GPT-5.6 with limited resources.
  • The biggest gap: What is still lacking is the "judgment” or ability to identify patterns from small amounts of data, a skill typically possessed by human researchers. While large models rely on massive datasets, scientific breakthroughs often require insights that cannot be derived from such data.

RSI is not an either-or situation; it’s a progression with multiple stages: the first step is improving efficiency (already achieved), the next is uncovering deeper patterns (underway), and the ultimate goal is for AI to make breakthroughs like Einstein did, using just a small amount of information.

V. Two underappreciated factors: Security and organization

  • Security: There’s less concern about AI getting out of control; however, practitioners believe this will take time. Training AI currently resembles training a dog, with reward and punishment mechanisms in place to prevent self-awareness. Anthropic has even suggested that coordinated global efforts to slow down development could be beneficial, but the challenge lies in how to monitor this process effectively (since it’s much more difficult than monitoring missile launch sites).
  • Organization: Large companies may not outperform smaller teams. RSI requires fast feedback, but large organizations have multiple layers of management, which can slow down information flow and revert to slower cycles. Smaller teams (with around 150 members, following Dunbar’s number) communicate more efficiently. Therefore, organizational flexibility might be as important as computational power.

VI. Which future do you believe in?

The potential of RSI is that the strong will get stronger, but there’s also the possibility of breakthroughs that could change the entire landscape. The allure of this competition lies in its uncertainty: you could predict a dominance by giants or a comeback by smaller teams. Regardless, RSI is an inevitable path for AI development. When AI can optimize itself, will humans lose control of its evolution? There’s no definitive answer, but it’s worth closely monitoring this trend.