ChatGPT users spent the week comparing notes after a sudden shift in performance: answers appeared sharper, more complete, and more capable, while response times stretched far beyond the usual pace. The pattern showed up across developers, creators, and benchmark watchers on X, leading to a wave of speculation that OpenAI may be running a quiet A/B test for GPT-5.6 and serving it to some Pro users under the GPT-5.5 Pro label.
Developer Anshu Chimala posted a side-by-side video showing differences in one-click landing page generation and joked that he had been lucky enough to get early access to GPT-5.6 Pro. Another developer, Dobroslav Radosavljevič, said the model he was using in Codex felt completely unlike 5.5. The replies split quickly. Some users treated the change as obvious proof of a model swap, while others argued the evidence was still thin.
Longer runs became the biggest clue
The strongest common signal was not a visible UI change but the amount of time the model was taking to finish hard tasks. Developer Conor Dart said a single prompt used to generate a 3D browser game with a physics engine and camera controls took more than one hour. He said GPT-5.5 Pro would normally finish similar work in about 10 minutes. His conclusion was simple: the result was imperfect, but the level reached from one prompt was still impressive.
AI commentator Chetas Lua described a similar pattern in robotics simulation tests, where response times stretched to roughly 20 to 40 minutes. He said he had not seen that kind of cadence since GPT-5.5 launched and claimed GPT-5.6 Pro was consistently beating Anthropic’s Fable 5 in 3D tasks.
Not every test pointed in the same direction. Benchmarker Chris compared two models using the same spaceship-building prompt and found that the suspected GPT-5.6 Pro ran for 87 minutes, while GPT-5.5 Extra High finished in 34 minutes and 42 seconds. His read was more restrained: GPT-5.6 looked like a steady, incremental upgrade over 5.5 rather than a decisive “Fable killer,” with mixed outcomes likely across different benchmarks.
Leaked specs focused on reasoning strength and recency
As the discussion spread, unverified technical details started circulating across multiple accounts. According to claims collected by leaker Pankaj Kumar, the model under test may use a knowledge cutoff pushed to December 2025. A reasoning-related setting referred to by some testers as “Juice Value” was also said to have increased from 768 to 960. The same posts claimed SVG and 3D design generation outperformed Fable 5 in certain tasks.
None of that has been confirmed by OpenAI. Even so, the recurring descriptions were strikingly similar from one account to another: stronger reasoning, an unfinished front-end experience, and a candidate build reportedly carrying the codename “Kindle-Alpha.”
AI commentator Leo, citing an anonymous message, said GPT-5.6 was being quietly tested on a subset of Pro accounts and that users selecting GPT-5.5 Pro might actually be running 5.6 in the background. He also pointed to June 25 as a possible public release date. That timeline remains speculative.
OpenAI has stayed silent before
The report notes that this would not be the first time OpenAI used a low-profile rollout. GPT-4.5 was also described as having been swapped in without advance notice, with confirmation coming only after users noticed differences in behavior. That kind of release style lets a company gather real usage data early and roll back quietly if the upgrade causes trouble.
OpenAI’s approach contrasts with Anthropic’s more explicit launch timelines for model generations such as Fable 5 and Mythos 5. The source material also says Chief Scientist Jakub Pachocki reportedly described the new model in an internal meeting as a “meaningful improvement” over GPT-5.5, while reporting from The Information stopped short of confirming any live A/B test or release schedule. Decrypt asked OpenAI about the rumors and had not received a response by publication time.
Competition is tightening around the model race
If OpenAI is moving faster, the competitive backdrop helps explain why. The report says China’s open-source GLM-5.2 trailed Claude Opus 4.8 by only 1 point on the FrontierSWE benchmark and had already moved ahead of GPT-5.5. That benchmark measures AI agent performance on complex engineering tasks that can run for hours, and it is gaining attention as a test of practical capability rather than short-form demos.
Anthropic is also dealing with pressure from another direction. Its flagship Mythos 5 and Fable 5 models were reportedly taken down after a U.S. export-control order issued on June 12, tied to a disputed jailbreak vulnerability. According to the report, that created a temporary gap in the top-end model market. A GPT-5.6 release before Anthropic resolves the issue could give OpenAI a useful opening.
The Wall Street Journal separately reported that OpenAI is evaluating price cuts for developers and enterprise customers as it prepares for a dual IPO. Against that backdrop, prediction activity has intensified. On Polymarket, the contract for a GPT-5.6 release between June 22 and June 28 had risen to 89% by the weekend. For now, though, the rumors remain just that until OpenAI confirms whether GPT-5.6 is real and whether it is already in users’ hands.

