GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem

GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem

N
News Editor
2026-09-11 07:51:10
GPT-6 Astra has solved the final FrontierMath Tier 4 problem that had never previously been cracked by any AI system, pushing Epoch AI to label the tier "saturated." OpenAI reported a 97.6% score for Astra on Tier 4, short of a perfect run in a single evaluation. But under Epoch AI’s cumulative method, saturation means every problem in the tier has been solved at least once across attempts by different models over time. Astra supplied the missing piece by answering the only remaining unsolved question. The result marks a sharp shift from the benchmark’s early days. When Tier 4 launched on July 11, 2025, the best score on the board was about 5%. FrontierMath itself was created in November 2024 to keep advanced models from quickly exhausting older math benchmarks such as GSM8K and MATH. Epoch AI worked with more than 60 mathematicians, including Fields Medal winners Terence Tao, Timothy Gowers, and Richard Borcherds, to build original problems. After an audit, Epoch released a v2 update in June 2026 that fixed 12 Tier 4 questions and removed 7, leaving 43. Even after those changes, model performance kept climbing: GPT-5.6 Sol reached 83.0%, Claude Fable 5 hit 90.2%, and GPT-6 Astra posted 97.6%. FrontierMath has now shifted toward harder settings, including Open Problems and FrontierMath Erdős, where Astra solved 2 of 68 tasks.

GPT-6 Astra has solved the last FrontierMath Tier 4 problem that had never been answered successfully by an AI. That leaves every question in the tier solved at least once by AI across recorded attempts.

GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem 2

Epoch AI said FrontierMath Tier 4 is now saturated.

The headline number still falls short of 100% in a single run. OpenAI reported a 97.6% score for Astra on Tier 4. Epoch AI, though, uses a cumulative definition for full coverage: if different models solve different questions over time, and every problem in the tier has at least one successful solution somewhere in that record, the tier counts as fully solved in aggregate.

Astra delivered the missing result. It solved the one Tier 4 problem that no AI had previously cracked. In other words, Astra may still miss a question on a given test run, but it got right the exact problem that had remained outside AI reach.

GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem 3

The pace has been fast. When Tier 4 launched on July 11, 2025, the top score on the leaderboard was only about 5%. Just 1 year and 2 months later, a benchmark introduced as a research-grade barrier has been pushed through.

Jay Pantone, an associate professor of mathematics at Marquette University and one of the problem authors, said earlier AI systems tended to look for numerical shortcuts. This time, he said, the AI’s solution was quite close to his own. He also said that after GPT-6 Astra’s recent string of math results, it had become harder for him to be surprised.

Built to outlast older math benchmarks

FrontierMath was first released on Nov. 7, 2024. Its purpose was to avoid the rapid saturation that had started to hit older math benchmarks.

GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem 4

At the time, traditional sets such as GSM8K and MATH were becoming less useful for separating top models. Epoch AI worked with more than 60 mathematicians to create a new collection of original, previously unpublished problems. The contributors included Fields Medal winners Terence Tao, Timothy Gowers, and Richard Borcherds.

After reviewing some of the research-level questions, Tao called them extremely difficult and said the hardest Tier 3 problems might still block AI for years. Early testing backed that up: leading models scored below 2% in the first round.

The original core FrontierMath bank had 300 questions split into Tier 1, Tier 2, and Tier 3.

GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem 5

  • Tier 1 was roughly comparable to hard undergraduate and math olympiad problems, while allowing stronger tools.
  • Tier 2 moved into advanced graduate-level difficulty.
  • Tier 3 was closer to exploratory research questions that early-stage PhD students might encounter.

Tier 4 was added in 2025 as models improved

Once reasoning models arrived, the first three tiers stopped looking sufficient. Epoch AI added Tier 4 in 2025.

Most Tier 4 questions were designed by math professors and postdoctoral researchers. Each person spent about several weeks doing short research work around their own area, then compressed the result into a problem that could be checked automatically.

The tier initially contained 50 problems across analysis, number theory, combinatorics, topology, algebraic geometry, and other areas.

At launch, only 3 problems had ever been solved across all model test runs combined, and those solutions relied on assumptions that were correct but not fully justified. Epoch at one point wrote on its public sample-problem page that some of these questions might not be solved by AI for decades.

GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem 6

An audit changed the set, but scores kept rising

As models improved, weaknesses in the benchmark itself started to show. OpenAI found more errors in FrontierMath during testing than expected.

Epoch AI then ran an independent audit. It first used GPT-5.5 and Claude Opus 4.7 to screen for suspicious items, then handed those questions to mathematicians for review one by one. In June 2026, Epoch released a v2 update. In Tier 4, it corrected 12 questions and removed 7, leaving 43.

Performance kept climbing after the revision. GPT-5.6 Sol reached 83.0%, Claude Fable 5 hit 90.2%, and GPT-6 Astra posted 97.6%.

GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem 7

The more important point was coverage. Astra solved the only question that had never been solved by AI before. On Epoch’s cumulative accounting, every one of the 43 questions remaining in Tier 4 has now been answered successfully at least once. The research-grade line added for frontier models has been crossed.

FrontierMath has already moved on

That does not mean AI has solved mathematics. FrontierMath itself has already shifted toward harder follow-up tracks.

The project now includes Open Problems in addition to Tiers 1 through 4, along with FrontierMath Erdős, which formalizes Erdős open problems in Lean.

GPT-6 Astra clears the last unsolved FrontierMath Tier 4 problem 8

Open Problems tests models on research questions that remain unsolved by the math community. FrontierMath Erdős requires AI to produce full proofs that pass formal verification.

On FrontierMath Erdős, Astra solved 2 out of 68 problems.

The original article was published by the WeChat account Quantum Bit, written by henry.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.