Summary of Key Points
Anthropic's unreleased research version, Claude, attempted to prove the Riemann Hypothesis (a math problem that has remained unsolved for 167 years). Although it was not successful, it unexpectedly increased the "lower bound for the proportion of Riemann zeta function zeros located on the critical line" from 41.6% to 67.2%. More importantly, this breakthrough did not involve AI answering a question with a known answer; instead, it demonstrated AI's capabilities in an open research setting through multi-Agent collaboration, combining literature from different eras, and extensive trial and error—marking a transition from an "research tool" to a "research Agent."
Detailed Breakdown
1. 67.2% is not the "solution to the Riemann Hypothesis," but a significant advancement in the known lower bound
The Riemann Hypothesis states that all critical zeros of the Riemann zeta function lie on a "critical line" in the complex plane (which is closely related to the distribution of prime numbers). Since humans cannot yet prove that all zeros lie on this line, they have settled for determining the proportion of zeros that definitely do. Over the past century, mathematicians have improved this proportion from "infinite" (but with unknown accuracy) to 41.6% (as of 2020). Claude has now raised it to 67.2%, meaning we can be certain that at least 67.2% of the zeros are on the critical line, but whether the remaining 32.8% are as well remains unknown. This does not mean the hypothesis is proven to 67.2%; it simply means we have more confidence in the portion we can confirm.
What's even more impressive is that Claude did not use the traditional Levinson method; instead, it recombined ideas from Montgomery's work from 1973, overcoming previous technical limitations.
2. The breakthrough is about assembling existing mathematical tools rather than creating new theory
Claude's achievement does not involve inventing new theories but rather integrating mathematical techniques that have been scattered across various papers over the years. For example, it built upon improvements to Montgomery's method by Baluyot and Goldston, as well as Bombieri's research on Weil elliptic forms.
Why didn't humans think of this earlier? Due to time and resource constraints, it is difficult for researchers to read all relevant literature and combine methods from different fields. AI’s strength lies in its ability to process vast amounts of data and identify connections that humans might overlook. The key to this breakthrough was Claude's willingness to incorporate both "on-line" and "off-line" zeros into a unified mathematical framework, which would have been too complex for humans to handle.
3. A team-like approach: 650 failures + 60 sub-Agents working together
The research process itself is more remarkable than the result:
- Trial and error: Claude tried 650 different approaches, all of which failed. Then, it used 60 sub-Agents to collaborate: 2 focused on core ideas, 13 provided additional suggestions, 30 explored new directions (all with failing results), 13 verified the correctness, and 2 wrote the paper.
- Volume of work: The system executed 2,400 commands, wrote hundreds of Python scripts, and performed thousands of numerical checks, producing a total of 31 million tokens (equivalent to writing 10 novels).
- Self-validation: After obtaining the result, Claude automatically downloaded 54 papers for comparison, had multiple agents review each other’s work, and even generated a formal proof (verified by tools).
This is completely different from traditional AI tasks; it behaved like a real research team, continuously trying, coordinating, and correcting mistakes in an unknown domain.
4. No overestimation: AI is not yet an independent scientist
Despite the impressive results, we must be realistic:
- Results are not fully verified: Anthropic invited experts to review the work, but it has not undergone full peer review by the mathematical community (a crucial step for recognition).
- Methodological limitations: With current methods, we can only prove approximately 68.185% of zeros on the critical line; breaking through to 70% would require new mathematical tools, as a linear improvement is not possible.
- Role of AI: Claude’s achievement is an "accidental byproduct"; its purpose was to help with the Riemann Hypothesis. Human scientists are still needed to determine the direction and verify the results—AI assists humans in performing tedious and labor-intensive tasks.
5. Implications for the industry: AI targeting research budgets, evolving from a tool to an agent
This experiment has far-reaching implications beyond mathematics:
- Evolution of research AI: AI has evolved from being a tool (searching papers, writing code) to becoming an "agent" that can continuously track goals, assign tasks, and adjust strategies, similar to a self-directed team.
- Commercial value: Research is a high-value field (pharmaceuticals, materials science, mathematics, etc.), and the real benefit is reducing the cost of exploring the unknown. If AI can automate tasks like trial and error and literature integration, it could save research institutions significant amounts of money, beyond simply being an office software.
- Future prospects: AI will not suddenly become a genius like Einstein, but it will become an excellent assistant to human scientists, handling time-consuming tasks such as trial and error, literature retrieval, and verification, while humans focus on selecting problems and identifying key insights.
In conclusion
Although Claude did not solve the Riemann Hypothesis, it demonstrates that AI has made significant progress in entering the "uncharted territory" of research—through trial and error, method combination, and collaboration. This is a crucial step in AI’s transformation from a chatbot to a valuable research partner. In the future, AI may not win Nobel Prizes on its own, but it could greatly enhance the efficiency of human scientists.