Tracking AI’s latest offline IQ test put several versions of GPT-5.6 at 136, according to the article. That makes GPT-5.6 the first large language model to move past the 130 mark on this benchmark. In the report’s framing, 130 is the lower boundary of the “genius” band in human IQ distribution, a level reached by about 1% of people worldwide.
136 on the offline set, not the public one
The article says Tracking AI uses two sets of questions. One is a public Mensa Norway-style test that anyone can take online, where models have already posted scores above 140. The other is an offline question bank assembled by Tracking AI itself. It is not public and is meant to reduce the risk that a model has already seen the answers.
The 136 result for GPT-5.6 came from that offline set. The report describes it as the harder and more cheat-resistant benchmark. On Tracking AI’s offline leaderboard, a range of GPT-5.6 variants, including a vision model, all reached 136 and opened a visible lead over the rest of the field.

Claude-5 Fable is listed next at 130. Below that, names including GPT-5.6 LUNA Max and Claude-4.8 Opus are described as landing between 117 and 123. The article adds that over the past year, models from o3 to other flagship systems had approached the 130 line but had not crossed it on this test. GPT-5.6, it says, is the first one through.
A whole model family reached the same score
The report emphasizes that this was not a one-off result from a single configuration. It says the SOL and TERRA branches of GPT-5.6 all hit 136, and the vision version kept pace as well.

It also cites a Reddit developer who tested the model directly and came away with the impression that GPT-5.6 felt clearly smarter than GPT-5.5. In the examples shown in the article, GPT-5.6 completed the test items quickly and posted strong results.
Developer trials: simulation, RAG workflow, and debugging
The article then shifts from benchmark scores to practical use. Developer Amir Bohlooli reportedly gave the same physics-simulation prompt to Fable 5 and GPT-5.6 Sol. He expected Fable to dominate, but the write-up says GPT-5.6 Sol was the more impressive result.

According to the article, GPT-5.6 Sol chose a particle fluid simulation, advanced the physics in real time rather than by running a fixed calculation on every frame, packed the CSS, interface, and rendering into a single HTML file, and then automatically hosted it as a shareable web page. One prompt produced a finished output.
Ramanpal Singh is cited as another example. Using a single prompt, he built a customer support ticketing system based on retrieval-augmented generation, or RAG. The article says the app included four roles, an admin backend, embeddable components, and features for classifying complaints, identifying sentiment, and drafting replies. It adds that he created five such apps and did so at a fraction of Fable 5’s cost.

The most vivid anecdote in the piece comes from Claire Vo. A few days earlier, she had been stuck on a bug and thought her own code might be broken. After switching to GPT-5.6 Sol, she left it with one line: “我就是不信搞不定.” The article says Sol fixed the issue in one pass and also helped other models run successfully afterward.
Her takeaway, as quoted in the report, is that Fable’s pursuit of technical perfection can trap it, while Sol’s more pragmatic style gets the job done.

Why the score does not settle the AGI question
The article notes that some users have already framed the result in stronger terms, with one online comment saying, “对99%的人来说,这已经是AGI了”. It then pulls back from that claim. The 136 score came from a specific offline benchmark by Tracking AI and from a Mensa Norway-style testing format. What it mainly measures, the report says, is standardized cognition: abstract pattern recognition and logical reasoning.
That leaves out a lot. IQ tests were not designed for large language models. A Mensa-style paper cannot tell readers how factually reliable a model is, how well it uses tools, or how dependable it will be in real professional settings. The article’s point is narrower: the test captures one slice of intelligence and shows how bright that slice is.

Still, the real-world examples matter in the way the article presents them. They suggest GPT-5.6 may be starting to bring “being good at tests” and “being good at work” closer together. The piece ends on that distinction. Standardized problems may resemble material a model has seen many times during training; the harder challenge is a genuinely new problem with no ready-made answer to copy.
Sources cited in the article
The article lists a post by davidpattersonx on X and the Tracking AI website as reference materials. It says the piece originally came from the WeChat public account “新智元” and credits the author as ASI启示录.

