Anthropic said this month that an unreleased research version of its Claude model pushed forward one of the oldest open problems in mathematics. The Claude Riemann bound, a guaranteed floor for the share of Riemann zeta zeros known to sit on the critical line, moved from 41.6% to 67.2% in a result the company published on its research site on October 5, 2026. The proof was formalized in the Lean proof assistant, which checks every logical step by machine, and reviewed by two of the company's mathematicians before publication. On a figure that mathematicians usually nudge upward in small increments over decades, a jump of more than twenty-five percentage points counts as a genuine advance.
What the bound actually measures
The Riemann hypothesis dates to 1859, when Bernhard Riemann conjectured that every nontrivial zero of the zeta function lies on a single line in the complex plane, the critical line. Those zeros encode information about how prime numbers are distributed, which is why the problem matters far beyond pure mathematics. It is one of the seven Millennium Prize Problems, each carrying a $1 million award for a proof, and no one has come close to proving it in more than a century and a half. Mathematicians instead settle for lower bounds, guarantees about a share of zeros known to behave as predicted. Anthropic did not claim to solve the hypothesis, and its own writeup says plainly that it does not expect this approach to do so.
According to the company's research writeup, reported on by TechSpot, Claude began by taking a real run at the full hypothesis and failed. It then turned to the bound problem, testing roughly 650 ideas that went nowhere before landing on a workable thread. The decisive run coordinated about 60 subagent instances over roughly a day and a half, executing 2,400 shell commands and generating around thirty-one million output tokens. The winning insight stitched together techniques from two previously published papers in a combination nobody had tried. Human researchers offered little direction beyond encouragement. Scientific American reports that the humans told the model to believe in itself and keep going.
Checked by machine, reviewed by people
Because the proof is formalized in Lean 4, its correctness is machine-checked rather than taken on trust. A proof written in Lean compiles or it does not, which leaves little room for hand-waving. Anthropic says two in-house mathematicians reviewed the result and outside experts were brought in before publication. Mathematicians quoted in the coverage treat the Claude Riemann bound as real research. Andrew Sutherland of MIT called it further evidence that AI can do interesting mathematical work instead of only solving problems handed to it. James Maynard of Oxford praised the writeup as restrained, noting that it avoided hype and credited earlier papers. Even being very optimistic, Maynard said, there is no pathway for this approach to settle the actual hypothesis.
The gap matters. The hypothesis demands every one of the zeros, and the new figure is a floor, not a finish line. Scientific American notes that even a perfect score on this related measure would not make the model eligible for the prize money, since the award goes to a proof of the hypothesis itself. Anthropic's statement likewise says it does not expect this technique to prove the full conjecture. What the Claude Riemann bound result shows instead is a documented case of an AI system producing a novel, verifiable mathematical result with minimal human guidance.
The Claude Riemann bound caps a busy autumn for AI mathematics. Late last month, Anthropic said a Claude model had produced the first complete machine-checked formalization of Fermat's Last Theorem, running to roughly thirteen million lines of Lean in under two weeks. Days later the company published a roundup of Claude-assisted work across eighteen scientific fields with nineteen human co-authors, including a solution to a lattice-integral problem first posed in 1939. Each announcement points in the same direction. The frontier labs have stopped arguing about which model scores best on coding tests and started racing to produce results that working mathematicians recognize as their own.
Comments 0
No comments yet. Be the first to share your thoughts!
Leave a comment
Share your thoughts. Your email will not be published.