GPT-5.6 Scores 136 in Tracking AI Offline IQ Test, Becoming the First LLM to Clear the 130 Mark

GPT-5.6 Scores 136 in Tracking AI Offline IQ Test, Becoming the First LLM to Clear the 130 Mark

N
News Editor
2026-07-16 08:21:08
Tracking AI’s latest offline IQ benchmark put several GPT-5.6 variants at 136, making it the first large language model family to move past the 130-point threshold on that specific test set. In the article, 130 is described as the point commonly associated with the lower edge of the “genius” range in human IQ distribution, a band reached by roughly 1% of people. The report says the result came not from a public Mensa-style test that models may already know, but from Tracking AI’s own offline, non-public question bank built to reduce answer leakage and memorization effects. Claude-5 Fable is listed at 130, while models such as GPT-5.6 LUNA Max and Claude-4.8 Opus are said to fall in the 117 to 123 range. The piece also cites developer anecdotes involving physics simulation, a RAG-based customer support ticketing system, and bug fixing, arguing that GPT-5.6 is showing strength both on standardized reasoning tasks and in practical work. At the same time, it notes that IQ tests capture only a slice of intelligence and do not measure factual reliability, tool use, or real-world job performance on their own.
GPT-5.6Tracking AIIQ testLLMClaudeAGIartificial intelligence

Tracking AI’s latest offline IQ test put several versions of GPT-5.6 at 136, according to the article. That makes GPT-5.6 the first large language model to move past the 130 mark on this benchmark. In the report’s framing, 130 is the lower boundary of the “genius” band in human IQ distribution, a level reached by about 1% of people worldwide.

136 on the offline set, not the public one

The article says Tracking AI uses two sets of questions. One is a public Mensa Norway-style test that anyone can take online, where models have already posted scores above 140. The other is an offline question bank assembled by Tracking AI itself. It is not public and is meant to reduce the risk that a model has already seen the answers.

The 136 result for GPT-5.6 came from that offline set. The report describes it as the harder and more cheat-resistant benchmark. On Tracking AI’s offline leaderboard, a range of GPT-5.6 variants, including a vision model, all reached 136 and opened a visible lead over the rest of the field.

GPT-5.6 Scores 136 in Tracking AI Offline IQ Test, Becoming the First LLM to Clear the 130 Mark 3

Claude-5 Fable is listed next at 130. Below that, names including GPT-5.6 LUNA Max and Claude-4.8 Opus are described as landing between 117 and 123. The article adds that over the past year, models from o3 to other flagship systems had approached the 130 line but had not crossed it on this test. GPT-5.6, it says, is the first one through.

A whole model family reached the same score

The report emphasizes that this was not a one-off result from a single configuration. It says the SOL and TERRA branches of GPT-5.6 all hit 136, and the vision version kept pace as well.

GPT-5.6 Scores 136 in Tracking AI Offline IQ Test, Becoming the First LLM to Clear the 130 Mark 4

It also cites a Reddit developer who tested the model directly and came away with the impression that GPT-5.6 felt clearly smarter than GPT-5.5. In the examples shown in the article, GPT-5.6 completed the test items quickly and posted strong results.

Developer trials: simulation, RAG workflow, and debugging

The article then shifts from benchmark scores to practical use. Developer Amir Bohlooli reportedly gave the same physics-simulation prompt to Fable 5 and GPT-5.6 Sol. He expected Fable to dominate, but the write-up says GPT-5.6 Sol was the more impressive result.

GPT-5.6 Scores 136 in Tracking AI Offline IQ Test, Becoming the First LLM to Clear the 130 Mark 5

According to the article, GPT-5.6 Sol chose a particle fluid simulation, advanced the physics in real time rather than by running a fixed calculation on every frame, packed the CSS, interface, and rendering into a single HTML file, and then automatically hosted it as a shareable web page. One prompt produced a finished output.

Ramanpal Singh is cited as another example. Using a single prompt, he built a customer support ticketing system based on retrieval-augmented generation, or RAG. The article says the app included four roles, an admin backend, embeddable components, and features for classifying complaints, identifying sentiment, and drafting replies. It adds that he created five such apps and did so at a fraction of Fable 5’s cost.

GPT-5.6 Scores 136 in Tracking AI Offline IQ Test, Becoming the First LLM to Clear the 130 Mark 6

The most vivid anecdote in the piece comes from Claire Vo. A few days earlier, she had been stuck on a bug and thought her own code might be broken. After switching to GPT-5.6 Sol, she left it with one line: “我就是不信搞不定.” The article says Sol fixed the issue in one pass and also helped other models run successfully afterward.

Her takeaway, as quoted in the report, is that Fable’s pursuit of technical perfection can trap it, while Sol’s more pragmatic style gets the job done.

GPT-5.6 Scores 136 in Tracking AI Offline IQ Test, Becoming the First LLM to Clear the 130 Mark 7

Why the score does not settle the AGI question

The article notes that some users have already framed the result in stronger terms, with one online comment saying, “对99%的人来说,这已经是AGI了”. It then pulls back from that claim. The 136 score came from a specific offline benchmark by Tracking AI and from a Mensa Norway-style testing format. What it mainly measures, the report says, is standardized cognition: abstract pattern recognition and logical reasoning.

That leaves out a lot. IQ tests were not designed for large language models. A Mensa-style paper cannot tell readers how factually reliable a model is, how well it uses tools, or how dependable it will be in real professional settings. The article’s point is narrower: the test captures one slice of intelligence and shows how bright that slice is.

GPT-5.6 Scores 136 in Tracking AI Offline IQ Test, Becoming the First LLM to Clear the 130 Mark 8

Still, the real-world examples matter in the way the article presents them. They suggest GPT-5.6 may be starting to bring “being good at tests” and “being good at work” closer together. The piece ends on that distinction. Standardized problems may resemble material a model has seen many times during training; the harder challenge is a genuinely new problem with no ready-made answer to copy.

Sources cited in the article

The article lists a post by davidpattersonx on X and the Tracking AI website as reference materials. It says the piece originally came from the WeChat public account “新智元” and credits the author as ASI启示录.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.