OpenAI published a public GitHub repository on the evening of October 6 containing 722 math manuscripts, every one generated by an internal model the company has not released. The papers are grouped into 372 result families, and the company says each family resolves or substantially advances an open question in mathematics or theoretical computer science. The claimed results include the quasi-Riemann hypothesis, the Kakeya problem in three and four dimensions, the Unique Games Conjecture, and a faster way to multiply very large integers. A detailed breakdown of the 372 result families shows how wide the release reaches.
Nothing on this scale has been attempted before. OpenAI says it posed the model roughly 4,000 open problems and kept only the output it judged significant. Each surviving result used an average of about three hours of ChatGPT Pro-equivalent compute, and the company says nearly all of them came from a single prompt handed to a single AI agent. That is a sharp drop from the company's September announcement about the Navier-Stokes problem, which Scientific American reported required a swarm of around 10,000 agents and millions of dollars of compute.
The reaction from working mathematicians has been immediate and split. Some welcomed the collection as a major gift to the field. Others said the company ignored guidance it had asked the community for, because the model and the prompts behind the results remain private. Within a day of publication, the repository's change history recorded the first corrections: three manuscripts were withdrawn after OpenAI found a sign error that invalidated an argument used across all three papers.
What the papers claim
The repository spans seventeen fields, from theoretical computer science and combinatorics to number theory, operator algebras, and mathematical logic. One family of papers claims a proof of the quasi-Riemann hypothesis: that every Dirichlet L-function, including the Riemann zeta function, has no zeros with real part greater than seven eighths. That would not settle the Riemann hypothesis, the most famous unsolved problem in mathematics, but researchers who have reviewed the claims say it would be real progress on it.
Other manuscripts claim a proof of Subhash Khot's Unique Games Conjecture, which concerns the limits of efficient approximation for optimization problems; results on the Kakeya problem in three and four dimensions; a proof that the irrationality exponent of pi is exactly two; and a deterministic integer-multiplication algorithm that would break a speed barrier believed optimal since 1971. The New York Times covered the release under the headline that OpenAI had released findings on nearly four hundred math problems, "further roiling" the field. The repository is published under an Apache 2.0 license, and corrections appear as new versions so that older versions stay accessible and citations keep working.
The verification problem
The strongest criticism of the math manuscripts is about verification rather than mathematics. OpenAI acknowledges the results sit at different stages of review. Some come with computer-checked proofs in Lean, a programming language that mechanically verifies every logical step. Many do not. According to the repository's own change history, 300 of the 719 current top-line results have Lean formalizations, roughly forty-two percent, and the company's README warns that unformalized results "could have issues." A passing Lean check confirms only that a proof follows from its stated premises, not that the result is new or important.
The three withdrawals show why the distinction matters. OpenAI said a sign error in a stabilization-trace cancellation argument corrupted a construction used by two dependent manuscripts, according to RuntimeWire's reading of the repository change log. The same history records revisions to fourteen other papers and reference updates in thirteen more. MIT mathematician Andrew Sutherland told Scientific American that until the model is released and the results are replicated, claims about solving problems with a single agent remain unverified, and that the field should "ask for receipts."
The release also fell short of the transparency standard the company had solicited. The Advisory Group on Mathematics and Artificial Intelligence, an independent panel of mathematicians based at the Institute for Advanced Study, published guidance on September 29 urging AI labs to release models, exact prompts, and compute figures for each result. OpenAI published the papers, ten abridged reasoning summaries, and average compute figures, but not the prompts, and the model remains internal. An OpenAI spokesperson told Scientific American that the company takes the guidelines seriously but is not bound by them, and that it is working to release the model as quickly and responsibly as possible.
Why it matters beyond mathematics
The dispute is the latest turn in a two-month collision between frontier AI labs and the mathematical community. In August, OpenAI met with around forty mathematicians to discuss what to do if AI outpaces humans in the field; WIRED reported that the group asked the company to publish real papers explaining the work rather than announcing it in blog posts. Northwestern mathematician Bryna Kra later told the magazine that, in her view, that input was ignored. The September Navier-Stokes claim drew further scrutiny after questions about related work by outside researchers, as reported by Axios.
The stakes reach past pure mathematics. Producing a proof is now cheap for whoever owns the model; checking it is still expensive for everyone else. Scientific American reported that OpenAI told the magazine many of the new results are not yet understood by the company's own mathematicians. If even a fraction of the result families survives scrutiny, the release would rank among the largest bursts of mathematical output ever recorded. The community will spend months finding out how much of it holds up.
The labs are also racing each other. Both OpenAI and its rival Anthropic are reportedly preparing to go public, and demonstrations of capability aimed at investors have become part of the mathematics story. The competition is playing out across the industry: Anthropic launched Claude Haiku 5.5 with a steep price cut this week, while the Genesis Mission is betting on superintelligence for scientific research. The Institute for Advanced Study warned in a statement that AI can now output mathematical arguments without the person who prompted them being able to understand the arguments, verify them, or take responsibility for them. For now, the checking is happening in public: the repository is open, the revision history is visible, and mathematicians have months of reading ahead.
Comments 0
No comments yet. Be the first to share your thoughts!
Leave a comment
Share your thoughts. Your email will not be published.