GPT-6 Astra posts first perfect score on Korean CSAT-style AI test, beating GPT-5.6 by 1.5 points

GPT-6 Astra posts first perfect score on Korean CSAT-style AI test, beating GPT-5.6 by 1.5 points

N
News Editor
2026-09-10 10:05:37
GPT-6 Astra scored a perfect 450 on a Korean College Scholastic Ability Test-style evaluation for large language models, becoming the first model to finish the full exam without losing a point. According to The Korea Daily, Astra earned full marks across Korean language, math, English, Korean history, and selected subjects including Physics I, Chemistry I, Biology I, and Social Studies and Culture. The runner-up, GPT-5.6, scored 448.5, leaving a gap of just 1.5 points. The score gap was narrow, but the compute gap was more pronounced. Astra used 357,000 tokens to complete the exam, versus 429,000 for GPT-5.6 and 562,000 for Claude Fable 5.1. That put Astra about 17% below the second-place model in token usage and 36% below Claude Fable 5.1. The test was run on GitHub using the actual 2026 academic year CSAT questions administered in 2025, with external web search barred during the evaluation. The report also cited professor Lee Seung-hyun, who said Astra’s edge appeared to come from structural efficiency in how it compressed and revised its solution plan while preserving its reasoning flow.

GPT-6 Astra scored a perfect 450 on a Korean College Scholastic Ability Test-style benchmark for AI models, becoming the first model to complete the full exam with no wrong answers.

According to The Korea Daily, Astra lost no points across Korean language, mathematics, English, Korean history, and elective subjects including Physics I, Chemistry I, Biology I, and Social Studies and Culture. On this evaluation, 450 is the maximum possible score and requires a completely error-free paper.

A tight ranking at the top

GPT-5.6 placed second with 448.5 points, just 1.5 behind Astra. GPT-5.4 ranked third at 448, Claude Fable 5.1 came fourth at 447.5, and Gemini 3.1 Pro was fifth at 445.

The top five models were separated by only 5 points. Measured against the 450-point total, the gap between first and second place came to about 0.3%.

The benchmark used a real Korean exam

The evaluation was run on a GitHub page titled the 2026 academic year CSAT LLM problem-solving record. It used the actual 2026 academic year Korean CSAT questions that were administered in 2025, meaning the models were tested on the same paper taken by real Korean students.

External web search was not allowed during the process. Models had to answer using only what they had learned during training. The report said that restriction made the scores more meaningful, because if online lookup were allowed, multiple-choice answers could already be found on the internet and the test would measure search ability rather than problem-solving ability.

The report also noted that test format can materially affect results. Gemini 3.1 Pro had previously been listed with a perfect 450 in an evaluation that covered only two subjects, but scored 445 in this broader, full-subject assessment. With more subjects and a wider spread of question types, scores became harder to maintain, which is why rankings need to be read alongside the number of subjects included in a benchmark.

Astra also led on token efficiency

Astra used 357,000 tokens to finish the entire paper. In this context, tokens are the units used to measure the text a model reads and writes. A higher token count generally means longer reasoning time and a more expensive run.

On the same exam, GPT-5.6 used 429,000 tokens and Claude Fable 5.1 used 562,000. Based on the figures in the report, Astra used about 17% fewer tokens than the second-place model and 36% fewer than Claude Fable 5.1, a difference of 205,000 tokens versus Claude.

That left an unusual split between score and compute use: only 1.5 points separated Astra from GPT-5.6, while the token gap was much larger.

Professor points to structural efficiency

The report cited professor Lee Seung-hyun, who said Astra appeared to maintain its reasoning flow while compressing and revising its solution plan. In his reading, the distinction was not simply raw intelligence but structural efficiency.

That interpretation matched the leaderboard shown in the report. Once scores are clustered near the ceiling, the place where models can still separate themselves is how many steps they need to solve the same question, and how many tokens they burn getting there.

Debate continues over what a perfect score means

On whether a perfect result on a college entrance-style test means artificial general intelligence has arrived, the original report said interviewed industry figures argued that exam scores are not the whole story and that AI’s value lies in solving problems that do not yet have answers. It also said OpenAI’s president and Jensen Huang have both stated that AGI has arrived, while debate continues in South Korean political and academic circles.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
7500

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.