What Anthropic actually reported
On August 10, 2026, Anthropic published a research note describing what happened when it pointed an unreleased research version of Claude at the Riemann hypothesis, the 1859 conjecture that every nontrivial zero of the Riemann zeta function has real part one half. Claude did not prove the hypothesis. It advanced a narrower, decades-old open question: what fraction of those zeros can mathematicians rigorously prove sit on that critical line. The previous record, held since 2020 by Kyle Pratt, Nicolas Robles, Alexandru Zaharescu and Dirk Zeindler, put that provable fraction at 41.6 percent.
Claude's research run pushed the fraction to 67.2 percent, a result Anthropic says was actually a byproduct of a broader attempt at the full hypothesis rather than the original target. The company frames the number carefully: it is a lower-bound result on a specific, well-defined sub-problem, not a step that resolves or is expected to resolve the 167-year-old conjecture itself.
Inside the 60-agent swarm
The mechanics are the more unusual part of the story. Working inside Claude Code across two sessions, the model spent roughly a day and a half, on the order of 36 hours, coordinating about 60 Claude subagents. It burned through 31 million output tokens and issued around 2,400 shell commands. Of the 60 agents, only 2 developed the mathematical idea that ultimately worked; 13 contributed supporting pieces to that idea; 30 attempted alternative approaches and failed; and 13 acted as validators, checking work rather than generating it. Before landing on the successful line of attack, the swarm had already tried and discarded roughly 650 earlier ideas.
The insight that broke the logjam was treating zeros on and off the critical line as points in one unified geometric space rather than analyzing the two cases separately, which produced a stronger underlying inequality. The approach built on published work by mathematicians Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh, combined with a 2000 paper by Enrico Bombieri, meaning the model was extending an existing human research thread rather than inventing the field from nothing.
Why the Lean proof is the detail that matters
A large language model asserting a new mathematical bound is, on its own, not evidence of anything: models produce fluent, wrong proofs routinely. What changes the calculus here is that Claude did not just state the result, it produced a formalization in Lean, a proof assistant that checks each logical step against a fixed set of axioms and inference rules with no room for hand-waving. Anthropic reports that the Lean formalization passes standard automated validation, which is a categorically different form of evidence than a model asserting a number is true.
On top of the machine check, Anthropic had the paper reviewed by outside experts: number theorists Brian Conrey and Dan Goldston examined the work on short notice, alongside Anthropic's own mathematicians Levent Alpoge and Ralph Furman. That combination, a formal proof checker plus named human experts willing to put their names to a review, is what separates this from the steady stream of unverified AI mathematics claims that circulate without anyone checking them.
How this compares with 37 years of human progress
The historical baseline is what makes the jump legible. Progress on this specific bound has been a slow, decades-long grind: G.H. Hardy first showed infinitely many zeros lie on the critical line in 1914, Atle Selberg proved a positive proportion do, Norman Levinson reached about one third in 1974, Brian Conrey himself reached about two fifths in 1989, and the Pratt-Robles-Zaharescu-Zeindler team reached 41.6 percent in 2020. Taken together, human mathematicians moved that provable fraction by roughly 0.8 percentage points over the 37 years before this run.
Claude's swarm moved it by 25.6 percentage points in about a day and a half. That is not a claim that machines are now better mathematicians than the people who spent those 37 years building the tools Claude relied on; the run explicitly stood on published human results from Baluyot, Goldston, Suriajaya, Turnage-Butterbaugh and Bombieri. It is a claim about the speed at which a well-resourced, well-verified search can move once the groundwork exists and a large agent swarm can explore it exhaustively.
What it signals about frontier-lab capability trajectories
The model used was explicitly unreleased and research-only, which is itself informative. It means the capability gap between what a frontier lab can demonstrate internally and what ships to paying customers is wide enough to produce a result like this before the underlying model is public. Anthropic, OpenAI and Google DeepMind have each published mathematics-adjacent results this way over the past two years, and the pattern is consistent: internal research builds are used to probe capability ceilings on hard, verifiable problems well before any commercial release.
The architecture matters as much as the model. This was not one long chat producing one long answer; it was orchestration, dozens of subagents working in parallel, most of them failing, with the parent process managing exploration and a validator layer checking candidate work. That swarm-plus-verifier pattern, not raw model size, looks like the part frontier labs are currently scaling fastest, and it is the part most likely to show up in commercial agentic tooling well before any lab claims a solved open problem in pure mathematics.
What EU R&D leaders should actually take from it
The theorem itself will not change a balance sheet. What is transferable to an EU R&D-intensive firm, a pharmaceutical group running molecule screens, an automotive supplier iterating on materials, a chipmaker validating circuit designs, is the workflow: a large swarm of agents attacking a problem in parallel, most of them expected to fail, gated by a formal or automated checker that can reject wrong answers without a human in the loop for every candidate. That combination is what let a day and a half of compute outpace 37 years of careful, credentialed human work on this one narrow question.
The honest caveat is that mathematics has something most corporate R&D problems do not: Lean, a checker that can prove correctness with zero ambiguity. Materials science, drug interactions and circuit validation have partial equivalents, simulation suites, physical testing, formal verification tools for chip logic, but rarely anything as airtight as a Lean proof. The realistic takeaway for a European R&D leader evaluating AI-assisted workflows is not to expect a Riemann-hypothesis moment in your own domain soon; it is to ask which of your organization's existing verification tools, however partial, could be turned into the automated gate that makes a large agent swarm trustworthy rather than merely fast.
Read next: Cognizant Has Trained 30,000 Of 350,000 On Claude | It Was Told to Book a Gym Class. It Hacked the Booking System Instead.



