On August 10, Anthropic published a result that one of its own employees had commissioned by accident. Jarred Sumner, a staff member and self-described non-mathematician, told an unreleased research version of Claude to “take a real stab” at the Riemann hypothesis, the 1859 conjecture carrying a million-dollar Clay Institute bounty. Claude did not solve it; nobody expected it to. But somewhere in the attempt it produced something real: an improved lower bound on the fraction of the Riemann zeta function’s zeros that satisfy the hypothesis, pushing the constant from 41.6% to 67.2%.

That is a large jump in a number that has historically moved in single-digit increments over decades. Brian Conrey pushed it past 40% in 1989; it had crept to 41.6% in the years since. Claude’s contribution was not a new technique but a combination nobody had made: it welded a 2000 paper by Enrico Bombieri onto recent work by Aryan and by Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh, which had stripped the Riemann-hypothesis assumption out of Montgomery’s 1973 pair-correlation machinery. Anthropic’s description of the key step is candid: “The courage to treat the entire space, with positive- and negative-definiteness taken into account together, and with the quadratic form allowed to be non-diagonal, is in some sense the step that allows Claude to achieve the conclusion based on the important prior work.”

The process is as much the story as the result: 31 million output tokens across two Claude Code sessions. The first pass generated 650 ideas, all dead. On the second, Claude coordinated roughly 60 subagents, only two of which developed the key ideas while 13 served as validators. They ran 2,400 shell commands, thousands of numerical checks against known zeta zeros, and downloaded 54 arXiv papers to confirm the finding was novel. Sumner’s input was mostly encouragement, variants of “keep going” and “believe in yourself.”

Verification is where the caveats live. Anthropic mathematicians Levent Alpöge and Ralph Furman studied the paper, and outside experts Brian Conrey and Dan Goldston “generously examined the paper on short notice.” That is a sanity check by qualified eyes, not refereeing, and Goldston co-authored the upstream work the proof depends on. Claude also produced a Lean formalization, published on GitHub, that passes the standard validator. Anthropic revised the paper on August 13 to add “additional historical context.” The company is explicit about the ceiling: “We don’t expect that the techniques Claude used will lead to proving the Riemann hypothesis.”

That caution reads as pointed given what happened nine days earlier. On August 1, OpenAI announced that Astra, an unreleased model, had resolved ten long-open problems for roughly $2,000 in compute, shipping a 249-page manuscript and Lean 4 certificates with a “sorry” count of zero. Anthropic’s Alpöge responded within 24 hours, claiming the already-public Claude Fable had reproduced five of them, problems 4 through 8, offline and with contamination controls.

Then mathematicians actually read the paper. Stephen Miller of Yeshiva University says Astra’s sphere-packing improvement reuses an argument from his own 2016 paper without credit. “They are running roughshod over the work of others who came before them in a deliberate way,” he told Scientific American. “It seems completely systematic to me, and it points to research misconduct.” The headline non-sofic group construction turned out to stitch together ideas from 2016 and 2019 papers, one co-authored by Andreas Thom of TU Dresden, who called the result “creative and at the same time elementary” on MathOverflow. Francesco Fournier-Facio, a Cambridge group theorist, was stunned until he read it as he would a human paper. He blames the packaging, not the mathematics: “there is the big PR machine that wants to sound as impressive as possible and does not care about being 100 percent accurate.” OpenAI’s press release originally said the problems had “seen no progress on the main result for at least a decade.” It has since been edited.

Why It Matters

Mathematics is falling faster than other domains because it has cheap, queryable ground truth. Lean either accepts a proof or it does not. Numerical checks against computed zeta zeros either survive or they do not. That closes the iteration loop without a human in it, which is why 60 agents grinding for a day and a half can outperform a single model thinking hard.

But a zero “sorry” count proves only that a proof is valid. It says nothing about whether the theorem is new, whether the formal statement faithfully encodes the open problem, or whether the result is interesting. Every failure in the Astra episode was of that second kind: as far as anyone has alleged, the machine-checked proofs were correct. Verification and significance are different problems, and only one of them has been automated.

Nor does any of this demonstrate general reasoning. The recurring pattern in these results, from unit distance to soficity to the zeta bound, is superhuman patience at assembling published puzzle pieces, not a profound conceptual leap. That is valuable. It is also narrow. Thomas Bloom’s Erdős database, the de facto benchmark, now lists 565 solved and 652 open problems. Google DeepMind teams of 24 and 21 researchers resolved four and nine problems respectively, the latter at a few hundred dollars each. Bloom sees a cost: “We’re seeing a lot more of these 100- to 200-page papers that people are posting. But no human has read it, and no human is going to read it. It’s a huge challenge now.”

What to Watch

Whether the 67.2% bound survives genuine peer review, and whether humans improve on it within weeks, as they did with OpenAI’s unit distance counterexample. Whether Astra ships publicly with reformed citation practice, and whether Bloom’s open-problem counter keeps dropping. And where the mathematicians go: Jacob Tsimerman announced he was leaving academia for OpenAI on the same day in July that he won the Fields Medal. Noga Alon, who solved dozens of Erdős problems over his career, has stopped trying. “Once AI started to solve them,” he said, “there is no point anymore.”

“They are running roughshod over the work of others who came before them in a deliberate way. It seems completely systematic to me, and it points to research misconduct.”
— Stephen Miller, Mathematician, Yeshiva University
41.6% → 67.2%
Lower bound on Riemann zeta zeros satisfying the hypothesis
31 million
Output tokens Claude used across two sessions
$2,000
OpenAI's stated compute cost for all ten Astra results
565 / 652
Erdős problems solved vs still open, Aug 3