Michael P. Frank, a researcher associated in the source text with reversible computing and computational physics, said he assigned GPT-5.6 Sol a demanding research task centered on Fermat’s Last Theorem. The goal was to examine whether there might be a path simpler than the Wiles-Taylor proof. The run continued for about 33 hours in the background, consumed significant computing resources, and was ultimately forced to stop by OpenAI’s system.
A long-running attempt focused on specific mathematical directions
According to Frank’s account, the assignment was not framed as an open-ended request to simply solve the theorem. It focused on several named directions: specialized modularity for Frey curves, uniform infinite descent, arithmetic abc-type inequalities, and uniform low-genus quotients. He also asked the model to keep rigorous notes, test candidate lemmas computationally, and clearly separate established results from conjectural ones.

No new proof emerged from the run. In its own retrospective analysis, GPT-5.6 Sol said there were two possible explanations for the forced stop. One was that the system judged the session to be using too many resources. The other was that OpenAI may already have tried similar work with related models and failed, making another extended run an unattractive use of compute. Frank appeared to share that assessment.
The model said it produced a map of dead ends, not a proof
GPT-5.6 Sol also gave an account of what it had done during those 33 hours. The substance of the work, it said, was to turn many ideas that sounded plausible into exact statements that could be checked, then disprove or eliminate them one after another. By its own description, the output was a map showing which routes did not work rather than a proof of the theorem itself. Its recommendation was to stop there.

Why the task was stopped is now part of a wider argument
That interruption quickly led to broader speculation. One line of interpretation held that OpenAI may prefer not to let a publicly accessible model solve an elite mathematical problem too casually, either because the company would not want outside users to seize the spotlight first or because it would rather keep that capability for internal use. In that reading, heavy dependence on a single platform could concentrate power over what discoveries become possible.
OpenAI senior research scientist Noam Brown pushed back on that idea. He said, “If someone just typed continue into our model, solved a Millennium Prize problem, and walked away with the $1 million prize, that would be the best possible advertisement for OpenAI. There’s nothing better than that.”

Critics of Brown’s argument answered that broad access to top-end capability might bring short-term publicity while still weakening OpenAI’s lead over time. In a competitive environment with multiple players, they argued, keeping stronger systems prioritized for internal research could matter more than letting outside users solve major problems first. Once a capability becomes widely available, in their view, the promotional value drops sharply.
Regulatory caution, routine safeguards, and a possible bug were all cited
The source text also described a regulatory angle. Under that interpretation, if a public-facing version of the model solved a very difficult mathematics problem, it would serve as open evidence that the model was extremely capable and might invite more regulatory pressure. In that context, some people viewed the idea of offering a weakened public model while retaining a stronger internal version as plausible.

Others pointed to a more technical explanation. In codex-cli, according to the account cited in the source, “goal blocked” means the model has run into the same blocker across 3 consecutive reasoning turns. If that is what happened here, the termination may have been part of an ordinary safety or anti-stall mechanism rather than a special restriction aimed at this specific theorem.
Another explanation was that the stop could have been caused by a bug. The discussion referenced in the source said Sol version 5.6 has an annoying bug during long-running tasks. A temporary workaround, it said, is to switch back to version 5.5, compress the context, let 5.5 run for a few minutes, and then switch back to Sol 5.6. The same discussion added that if a user wants to preserve the same session, it can get stuck and fail to come back out.

The episode fed a broader debate about AI and advanced mathematics
The report connected the incident to the recent rise in discussion around mathematics and AI following the announcement of the Fields Medal. It framed the underlying question in plain terms: if an ordinary person hands a hard math problem to an AI system and keeps telling it to continue, can that process eventually produce a genuine breakthrough, win a major prize, or alter the course of mathematical development?
At least in this case, the answer was no. After 33 hours, GPT-5.6 Sol had not produced a new proof of Fermat’s Last Theorem. What it left behind was a record of eliminations, along with a fresh argument over compute limits, public access to advanced models, regulatory pressure, and the practical details of long-running AI research sessions.

The source text also cited a reference link: https://x.com/lu_sichu/status/2081367506468360495. It said the original article came from the WeChat public account "机器之心," written by Zhang Qian.

