GPT-6 Astra has solved the last FrontierMath Tier 4 problem that had never been answered successfully by an AI. That leaves every question in the tier solved at least once by AI across recorded attempts.

Epoch AI said FrontierMath Tier 4 is now saturated.
The headline number still falls short of 100% in a single run. OpenAI reported a 97.6% score for Astra on Tier 4. Epoch AI, though, uses a cumulative definition for full coverage: if different models solve different questions over time, and every problem in the tier has at least one successful solution somewhere in that record, the tier counts as fully solved in aggregate.
Astra delivered the missing result. It solved the one Tier 4 problem that no AI had previously cracked. In other words, Astra may still miss a question on a given test run, but it got right the exact problem that had remained outside AI reach.

The pace has been fast. When Tier 4 launched on July 11, 2025, the top score on the leaderboard was only about 5%. Just 1 year and 2 months later, a benchmark introduced as a research-grade barrier has been pushed through.
Jay Pantone, an associate professor of mathematics at Marquette University and one of the problem authors, said earlier AI systems tended to look for numerical shortcuts. This time, he said, the AI’s solution was quite close to his own. He also said that after GPT-6 Astra’s recent string of math results, it had become harder for him to be surprised.
Built to outlast older math benchmarks
FrontierMath was first released on Nov. 7, 2024. Its purpose was to avoid the rapid saturation that had started to hit older math benchmarks.

At the time, traditional sets such as GSM8K and MATH were becoming less useful for separating top models. Epoch AI worked with more than 60 mathematicians to create a new collection of original, previously unpublished problems. The contributors included Fields Medal winners Terence Tao, Timothy Gowers, and Richard Borcherds.
After reviewing some of the research-level questions, Tao called them extremely difficult and said the hardest Tier 3 problems might still block AI for years. Early testing backed that up: leading models scored below 2% in the first round.
The original core FrontierMath bank had 300 questions split into Tier 1, Tier 2, and Tier 3.

- Tier 1 was roughly comparable to hard undergraduate and math olympiad problems, while allowing stronger tools.
- Tier 2 moved into advanced graduate-level difficulty.
- Tier 3 was closer to exploratory research questions that early-stage PhD students might encounter.
Tier 4 was added in 2025 as models improved
Once reasoning models arrived, the first three tiers stopped looking sufficient. Epoch AI added Tier 4 in 2025.
Most Tier 4 questions were designed by math professors and postdoctoral researchers. Each person spent about several weeks doing short research work around their own area, then compressed the result into a problem that could be checked automatically.
The tier initially contained 50 problems across analysis, number theory, combinatorics, topology, algebraic geometry, and other areas.
At launch, only 3 problems had ever been solved across all model test runs combined, and those solutions relied on assumptions that were correct but not fully justified. Epoch at one point wrote on its public sample-problem page that some of these questions might not be solved by AI for decades.

An audit changed the set, but scores kept rising
As models improved, weaknesses in the benchmark itself started to show. OpenAI found more errors in FrontierMath during testing than expected.
Epoch AI then ran an independent audit. It first used GPT-5.5 and Claude Opus 4.7 to screen for suspicious items, then handed those questions to mathematicians for review one by one. In June 2026, Epoch released a v2 update. In Tier 4, it corrected 12 questions and removed 7, leaving 43.
Performance kept climbing after the revision. GPT-5.6 Sol reached 83.0%, Claude Fable 5 hit 90.2%, and GPT-6 Astra posted 97.6%.

The more important point was coverage. Astra solved the only question that had never been solved by AI before. On Epoch’s cumulative accounting, every one of the 43 questions remaining in Tier 4 has now been answered successfully at least once. The research-grade line added for frontier models has been crossed.
FrontierMath has already moved on
That does not mean AI has solved mathematics. FrontierMath itself has already shifted toward harder follow-up tracks.
The project now includes Open Problems in addition to Tiers 1 through 4, along with FrontierMath Erdős, which formalizes Erdős open problems in Lean.

Open Problems tests models on research questions that remain unsolved by the math community. FrontierMath Erdős requires AI to produce full proofs that pass formal verification.
On FrontierMath Erdős, Astra solved 2 out of 68 problems.
The original article was published by the WeChat account Quantum Bit, written by henry.

