OpenAI on Tuesday published 722 mathematics manuscripts on GitHub, saying they were produced by an internal model the company has not released. An OpenAI spokesperson said almost everything came from a single prompt given to a single AI agent, though some results may have taken multiple attempts.
The claim is sweeping, and potentially important for mathematics. It has also met immediate skepticism from researchers who say the model is unavailable, the results are hard to reproduce, and much of the material still lacks full verification.
What OpenAI actually released
The 722 papers are grouped into 372 families of related results. A family can include a main theorem, companion arguments, consequences, or alternative proofs. That means 722 refers to manuscripts, not solved problems.
OpenAI said it posed roughly 4,000 problems to the model and kept the outputs it judged significant enough to publish.
The company said the average result used the equivalent of about three hours of ChatGPT Pro thinking compute. It contrasted with last month’s Navier-Stokes claim, which involved 10,000 coordinating agents running for 88 hours.
Only part of the set has Lean-checked results
OpenAI released abridged reasoning summaries for 10 results. Even so, only 162 of the 722 manuscripts currently come with a computer-checked main result, according to a formalization catalog in the repository. That works out to about 22% of the collection. Those results were translated into Lean, software that mechanically checks each logical step.
OpenAI said not all manuscripts have Lean formalizations and that 「some of the unformalized results could have issues.」 In plain terms, some of what the company published may be wrong.
A Lean pass also has limits. It can show that a proof follows from the statement as written in Lean, but it does not establish that the Lean statement matches the original problem, or that the result is new or important. Mathematicians still have to judge those points themselves.
Researchers ask for reproducibility and a public review path
Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology, told Scientific American: 「Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified.」 He added: 「We should ask for receipts.」
Some of the criticism has focused on the content of the proofs themselves. Dmitry Rybin wrote on X that he was trying to read OpenAI’s proof that the chromatic number of a plane is >= 6, but found it 「totally unbelievable alien math.」 He said the model somehow found that any K-coloring is equivalent to a 「weakly measurable」 K-coloring, which he said seemed to come out of nowhere.
Keith Adler raised a different issue. He wrote that the openai/math repository has Issues turned off and has never accepted a pull request. He called that disappointing, saying that if OpenAI publishes 722 manuscripts and asks for Lean formalizations, it should provide somewhere for people to submit them. Adler also said he is formalizing OpenAI’s Saxl’s Conjecture proof in Lean 4.
Institute for Advanced Study stresses human understanding
The Institute for Advanced Study in Princeton, New Jersey, said in a statement: 「It is now the case that AI can output mathematical arguments in situations without the human who prompted it being able to understand the arguments, verify them, or take responsibility for them.」 The institute added: 「We believe that human understanding of mathematics remains of paramount importance. How, in this new era, can we work towards a new paradigm that includes human understanding of mathematics as part of responsible scholarly output?」
The statement sharpened the debate around scholarly responsibility. If a result comes from reasoning that humans cannot readily interpret, then review, attribution, and accountability become harder to define.
Some mathematicians called it a major day for the field
Not every reaction was skeptical. Professor Abhishek Saha wrote: 「It is a very big day for mathematics,」 while also saying that most of the problems fit better under 「exceptional advances within an existing program」 than under 「surprising breakthroughs.」
Under that framing, most of the released work is interesting, but not in the category of the millennium problems. Saha said exactly one problem in the 722-manuscript set belongs in that top tier: the Quasi-Riemann Hypothesis.
In another post, Saha said he had further thoughts on the 372 results released by OpenAI across 722 manuscripts and outlined a rough classification system for theorems by how groundbreaking they are.
The release still falls short of outside recommendations
The publication also did not meet the standard recommended on September 29 by an advisory group at the Institute for Advanced Study. That group said each result should come with the model name, the prompts, a summarized chain of thought, the time taken, and the compute cost.
OpenAI published average compute figures and 10 reasoning summaries, but not the prompts. The company said it is still working on how to release the model responsibly.
Daniel Litt, a mathematician at the University of Toronto, took the opposite view from those urging restraint. He argued there is no reason to ask the company to keep the answers to these math questions secret.
A contrast with Anthropic’s earlier release
Anthropic took a different approach last month with its Lean-checked proof of Fermat’s Last Theorem. The company posted all 13 million lines publicly on GitHub. That work formalized a theorem Andrew Wiles published in 1995, rather than claiming new mathematical results.
OpenAI said it will add Lean formalizations as it obtains them. For now, 162 of the 722 manuscripts include one.

