GPT-6 Astra scored a perfect 450 on a Korean College Scholastic Ability Test-style benchmark for AI models, becoming the first model to complete the full exam with no wrong answers.
According to The Korea Daily, Astra lost no points across Korean language, mathematics, English, Korean history, and elective subjects including Physics I, Chemistry I, Biology I, and Social Studies and Culture. On this evaluation, 450 is the maximum possible score and requires a completely error-free paper.
A tight ranking at the top
GPT-5.6 placed second with 448.5 points, just 1.5 behind Astra. GPT-5.4 ranked third at 448, Claude Fable 5.1 came fourth at 447.5, and Gemini 3.1 Pro was fifth at 445.
The top five models were separated by only 5 points. Measured against the 450-point total, the gap between first and second place came to about 0.3%.
The benchmark used a real Korean exam
The evaluation was run on a GitHub page titled the 2026 academic year CSAT LLM problem-solving record. It used the actual 2026 academic year Korean CSAT questions that were administered in 2025, meaning the models were tested on the same paper taken by real Korean students.
External web search was not allowed during the process. Models had to answer using only what they had learned during training. The report said that restriction made the scores more meaningful, because if online lookup were allowed, multiple-choice answers could already be found on the internet and the test would measure search ability rather than problem-solving ability.
The report also noted that test format can materially affect results. Gemini 3.1 Pro had previously been listed with a perfect 450 in an evaluation that covered only two subjects, but scored 445 in this broader, full-subject assessment. With more subjects and a wider spread of question types, scores became harder to maintain, which is why rankings need to be read alongside the number of subjects included in a benchmark.
Astra also led on token efficiency
Astra used 357,000 tokens to finish the entire paper. In this context, tokens are the units used to measure the text a model reads and writes. A higher token count generally means longer reasoning time and a more expensive run.
On the same exam, GPT-5.6 used 429,000 tokens and Claude Fable 5.1 used 562,000. Based on the figures in the report, Astra used about 17% fewer tokens than the second-place model and 36% fewer than Claude Fable 5.1, a difference of 205,000 tokens versus Claude.
That left an unusual split between score and compute use: only 1.5 points separated Astra from GPT-5.6, while the token gap was much larger.
Professor points to structural efficiency
The report cited professor Lee Seung-hyun, who said Astra appeared to maintain its reasoning flow while compressing and revising its solution plan. In his reading, the distinction was not simply raw intelligence but structural efficiency.
That interpretation matched the leaderboard shown in the report. Once scores are clustered near the ceiling, the place where models can still separate themselves is how many steps they need to solve the same question, and how many tokens they burn getting there.
Debate continues over what a perfect score means
On whether a perfect result on a college entrance-style test means artificial general intelligence has arrived, the original report said interviewed industry figures argued that exam scores are not the whole story and that AI’s value lies in solving problems that do not yet have answers. It also said OpenAI’s president and Jensen Huang have both stated that AGI has arrived, while debate continues in South Korean political and academic circles.

