A debate over unpublished AI-generated mathematics has spread quickly through the research community after claims that major results may already be sitting inside private labs.

In a Sept. 15 blog post, Scott Aaronson, a theoretical computer science professor at the University of Texas at Austin and a former visiting researcher with OpenAI’s alignment team, wrote that his 13-year-old daughter joked, 「If I want to become a mathematician, it looks like I have about two weeks left.」 In the same post, Aaronson listed several long-standing open problems that AI had recently proved or helped prove, including a counterexample to the Jacobian conjecture, a Lean-verified version of Fermat’s Last Theorem, and long-pending questions in quantum complexity theory.
The most sensitive part of the post came at the end of that list. Aaronson said that after a hostile public response to work on the Navier-Stokes equations, AI companies were now keeping solutions to some 「very significant problems」 private while they figured out how to handle them.
Ethan Mollick, an associate professor at the Wharton School, reposted the claim and said he does not usually pass along rumors, but believed this one likely came from someone with inside information. Scott Armstrong, a mathematics professor at New York University, then added a sharper claim: he had been told that OpenAI had been sitting on 「hundreds」 of mathematical proofs since at least the International Congress of Mathematicians, or ICM, in July.

That has shifted the anxiety inside mathematics. The issue is no longer only whether AI can solve famous problems. It is whether researchers outside a handful of labs even know which problems may already be solved.
Proof production is now measured in hours
In Aaronson’s telling, the idea of hundreds of unpublished proofs sounds extreme, but it fits the operating model of a frontier AI lab.
On Aug. 1, OpenAI released 10 results in mathematics and theoretical computer science, saying each one either solved or materially advanced a long-standing open problem. On Sept. 1, according to the account, OpenAI heard that two Millennium Prize Problems had been solved. Not knowing which two, the company fed all remaining unsolved Millennium problems, along with a set of other high-impact questions, into an internal model described as more capable than GPT-6 Astra.
A team of nearly 100 agents reportedly spent about 50 hours solving the no-external-force version of the Euler equations. Resources were then redirected to Navier-Stokes. The group working on that problem used roughly 10,000 agents in parallel and produced a solution about 88 hours after the first agents were launched. Lean formal verification then took another 17 hours.

Aaronson estimated that the Navier-Stokes proof alone consumed about 130 billion output tokens. He also estimated that the 166-page proof burned at least $15 million in compute, while adding that no human appears to have fully understood it yet.
Armstrong said labs can now produce a 160-page paper within days, along with a Lean formal proof running to hundreds of thousands of lines. In that setting, mobilizing 10,000 agents for 88 straight hours after hearing a rumor is not a thought experiment. It is a description of what concentrated compute can do.
Machines produce results in days, humans absorb them in years
The pace of AI output and the pace of academic validation are no longer aligned.

Under Clay Mathematics Institute rules, a Millennium Prize Problem is not recognized as solved until the work is formally published, survives a two-year waiting period, and passes strict outside review. Clay still lists Navier-Stokes as unsolved and said it would examine the episode in detail.
That gap matters. A machine may generate a result in days. Human experts may need years to read it, test it, present it, and fold it into the literature. Peer review, conference talks, and textbook adoption were built for a much slower research cycle.
From that perspective, a backlog of unpublished proofs inside a lab is not an oddity. It is a predictable consequence of output arriving faster than the academic system can process it.
Disclosure, priority, review, and authorship are all under strain
The first norm under pressure is disclosure. Aaronson referred to 「AI companies」 in the plural, while Armstrong named OpenAI directly and attached the number 「hundreds」 to the claim. Armstrong, who is affiliated with Sorbonne University and the French National Centre for Scientific Research, or CNRS, said the concern is not only how many proofs may be sitting unpublished, but that access to knowledge may be shifting from public availability to insider awareness. He wrote that rumors about major results are circulating widely, and that a person’s social distance from a lab may determine how much they know.

The second norm is priority. Traditionally, the person who proves and publishes first gets the credit. Aaronson’s account says the Navier-Stokes episode disrupted that rule. New York University mathematician Tristan Buckmaster and Anthropic’s Levent Alpöge had already made progress in that direction. OpenAI, after hearing that work was underway, reportedly pushed massive resources into the problem in less than a week and later contacted the two researchers about releasing simultaneously.
Terence Tao warned on Mastodon that if word spreads that someone is working on a problem, large-scale AI compute may now descend on it before the human project has time to mature. That would give researchers a reason to stop sharing promising ideas with peers.
The third norm is the expectation that a result should be understood before publication. Lean can certify logical correctness, but it does not guarantee that humans can follow the proof quickly. Aaronson compared journal editors to the defenders of Gondor in The Lord of the Rings, trying to hold a wall against an overwhelming assault. His point was blunt: without handing some review work to AI, the human system cannot keep up.

On Sept. 11, Tao and 25 Fields Medalists published an open letter criticizing AI companies for treating hard mathematics as a benchmark target. The letter argued that mathematics is ultimately about understanding, and that solving a problem is only one part of that goal.
The fourth norm is authorship. If a mathematician barely participated in finding the proof, should that person still be listed as an author? If a model did most of the work, what would it even mean to put GPT-6 Astra in the author line? Aaronson said conversations among mathematicians now tend to circle back to AI no matter where they begin.
A new uncertainty now hangs over mathematical work
The deeper fear is easy to state. A doctoral student may have spent three years on a hard problem while the answer has already been generated inside an AI lab and never disclosed. The student does not know. The adviser does not know. Reviewers do not know.
That changes the structure of competition. Academia used to ask who would prove a theorem first. It may now also have to ask who knows which theorems have already been proved. Questions about timing, format, credit, and control are no longer abstract.

Aaronson wrote in his blog, 「The singularity has begun, but it is very unevenly distributed.」 He added that he still has to unload the dishwasher and clip his toenails, but that for the rest of his life he probably will not prove a theorem because he is truly needed. If he proves one, he said, it will be for fun.
What looks uneven here is not only problem-solving power. It is also access to information and the sense of what mathematical work is worth. For centuries, the central question was who could solve a problem. Now two more questions are creeping in: whether the problem you are working on has already been solved by a machine, and whether sharing your idea too early could invite enough compute to finish the job before you do.
Mathematics has long depended on open circulation. A proof is written down, others read it, test it, and build on it. If answers begin by landing on private servers inside AI labs, that tradition faces a direct challenge.
References
- Scott Aaronson blog: https://scottaaronson.blog/?p=10062
- Ethan Mollick on X: https://x.com/emollick/status/2099996799750623665

