An unreleased research version of Anthropic’s Claude AI has produced a significant new mathematical result while attempting to tackle the Riemann hypothesis, one of the most famous unsolved problems in mathematics.
The AI did not solve the Riemann hypothesis itself. Instead, during its attempt, Claude improved a longstanding mathematical bound connected to the problem, raising the known minimum proportion of certain zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
The result offers another indication that advanced AI systems may be developing the ability to contribute meaningfully to mathematics, even when they cannot fully solve the difficult problems they are given.
The Riemann hypothesis has challenged mathematicians for more than 150 years.
First proposed by mathematician Bernhard Riemann in 1859, the hypothesis concerns the Riemann zeta function and the distribution of prime numbers.
Prime numbers are numbers greater than one that can only be divided evenly by one and themselves. They form one of the foundations of number theory, but their distribution across the number line appears irregular.
The Riemann hypothesis proposes a precise mathematical pattern involving the zeros of the zeta function.
Proving that the pattern always holds would have major implications across number theory and several related areas of mathematics.
The problem is also one of the Clay Mathematics Institute's Millennium Prize Problems, with a $1 million prize available for a valid proof.
No mathematician or AI system has yet produced an accepted general proof.
Anthropic's model was originally asked to make a serious attempt at solving the Riemann hypothesis.
It failed to prove the hypothesis, but the process led to an unexpected result.
Mathematicians have already established that a certain proportion of the relevant zeros of the Riemann zeta function lie on the line predicted by the hypothesis.
Before Claude's work, the established lower bound was 41.6%.
Anthropic says its unreleased model developed an argument that increased that lower bound to 67.2%.
In practical terms, the result does not demonstrate that all of the zeros behave as predicted by the Riemann hypothesis. It does, however, substantially increase the proportion that can be shown to satisfy the relevant condition.
That distinction is important because the AI has not solved the original problem.
Instead, it appears to have made progress on a related mathematical question that researchers have studied for decades.
The result did not emerge completely independently of previous mathematical research.
Claude drew heavily from existing work by mathematicians, including research by several number theorists as well as earlier work by mathematician Enrico Bombieri.
The model appears to have found a new way to combine these established techniques.
Anthropic says the approach allowed Claude to move beyond the previous 41.6% lower bound and reach 67.2%.
This type of result could prove especially important for understanding how AI contributes to science and mathematics.
An AI system does not necessarily need to invent an entirely new field or solve a famous problem from scratch to make a useful contribution. Finding overlooked connections between existing ideas could itself become an important capability.
Claude's mathematical search was extensive.
An Anthropic staff member who was not a professional mathematician initially asked the model to make a genuine attempt at proving the Riemann hypothesis.
Claude generated and tested around 650 different ideas during its early attempts, but none produced a successful solution.
The process then expanded considerably.
The model spent roughly a day and a half coordinating around 60 AI subagents that explored different parts of the mathematical problem.
Together, the agents executed thousands of computational tasks and produced hundreds of Python scripts.
They also performed numerical checks using known zeros of the Riemann zeta function and reviewed one another's mathematical work.
The experiment also illustrates how computationally intensive AI-assisted mathematical research can become.
Anthropic says the model used approximately 31 million output tokens across two Claude Code sessions.
The group of AI agents performed different roles during the process.
A small number developed the main mathematical ideas, while others contributed possible approaches, attempted alternative methods, reviewed arguments, checked correctness, and helped prepare the resulting paper.
This multi-agent approach allowed the model to explore many possible directions while also assigning separate AI instances to verification.
Such systems could become increasingly important as researchers experiment with giving AI models longer-running scientific and mathematical tasks.
Producing a promising mathematical argument is only part of the challenge.
The argument must also survive rigorous verification.
After reaching its result, Claude used additional agents to review the proof, search for possible counterexamples, and independently reproduce parts of the reasoning.
The system also examined dozens of existing research papers to determine whether the result had already been discovered.
Claude eventually recommended that human mathematicians examine the work.
Two mathematicians working at Anthropic then reviewed and validated the result.
Outside experts in the field were also asked to examine the research.
Another important part of the process involved formal verification.
Claude's result was translated into Lean, an open source proof assistant used to formally verify mathematical arguments.
Proof assistants allow mathematical statements and logical steps to be expressed in a form that a computer can systematically check.
This provides an additional layer of verification beyond conventional review.
Formal proof tools are becoming increasingly significant in AI mathematics because they can help reduce one of the biggest risks associated with language models: producing reasoning that appears convincing but contains subtle logical errors.
Combining powerful AI models with formal verification systems could therefore become an important model for future AI-assisted research.
Anthropic has been careful to distinguish the new result from a proof of the Riemann hypothesis.
The company does not expect the techniques used by Claude in this work to directly produce a complete solution to the famous problem.
The Riemann hypothesis remains unsolved.
Instead, the experiment demonstrates that a modern AI model can be given an extremely difficult mathematical goal, explore a large number of possible approaches, coordinate multiple agents, study previous research, and potentially discover a useful result along the way.
That may be significant even when the original objective remains out of reach.
The Claude experiment is part of a broader trend involving AI systems and advanced mathematical research.
Language models were once primarily evaluated on whether they could answer known mathematics questions correctly.
Increasingly, researchers are testing whether models can contribute to problems where the answers are not already known.
The difference is substantial.
Solving an established textbook problem largely measures whether a model can reproduce or reconstruct existing knowledge.
Making progress on an open mathematical problem requires something closer to research. The system must explore uncertain ideas, discard unsuccessful approaches, connect different areas of prior work, and develop arguments that have not simply been copied from known solutions.
Despite the progress, human mathematicians continue to play a crucial role.
AI models can generate large amounts of mathematical reasoning, but sophisticated errors may be difficult to identify without expert review.
Independent verification is particularly important when models work on open research problems because there may be no existing answer against which the output can be compared.
In Claude's case, human mathematicians examined the result, while formal verification provided an additional check.
That combination of AI exploration, expert review, and machine-verifiable proofs could become a common structure for AI-assisted mathematics.
The experiment also raises broader questions about the future role of mathematicians.
AI agents capable of generating hundreds of ideas and investigating them simultaneously could dramatically expand the number of possible approaches researchers can explore.
Instead of manually testing every mathematical direction, a researcher could potentially ask AI agents to investigate many possibilities and then focus human attention on the most promising results.
This could turn AI into a powerful research collaborator rather than simply a tool for answering questions.
At the same time, the growing role of AI in mathematical discovery is creating debates around authorship, responsibility, verification, and how credit should be assigned when machines contribute substantially to a result.
The most notable aspect of Anthropic's experiment may be that the model discovered the result while unsuccessfully pursuing a much larger goal.
Claude was asked to attack the Riemann hypothesis. It did not solve it, but its search produced progress on a related question.
That pattern could become increasingly common in AI-assisted science.
A model may fail at an ambitious objective while still discovering useful techniques, intermediate results, unexpected connections, or new research directions along the way.
For mathematics, this could mean that the impact of advanced AI will not be measured only by whether a machine suddenly solves one of the world's most famous problems.
The more immediate transformation may come from AI systems becoming capable research partners that explore mathematical territory at a scale and speed that humans alone cannot easily match.
Share your thoughts about this article.
Be the first to post a comment!