Meta AI chief Alexandr Wang pushes back on SemiAnalysis claim that models were over-optimized for benchmarks

Meta AI chief Alexandr Wang pushes back on SemiAnalysis claim that models were over-optimized for benchmarks

N
News Editor
2026-09-08 08:22:17
Meta Chief AI Officer Alexandr Wang has rejected a SemiAnalysis argument that Gemini 3.8 Flash and Muse Spark 1.3 were heavily optimized for public benchmark performance. SemiAnalysis had pointed to the two models’ sharp score drop between Terminal-Bench 2.1 and the newer Terminal-Bench 4.0, saying the pattern made them some of the clearest examples of “benchmaxxed” systems. Gemini 3.8 Flash and Muse Spark 1.3 scored 89.4% and 88.8% on Terminal-Bench 2.1, close to GPT-6 Astra and Claude Fable 5.1, but fell to 19.1% and 33.3% on Terminal-Bench 4.0. Astra and Fable 5.1 scored 57.7% and 55.8% on the newer test. SemiAnalysis said public benchmarks may have become too easy to target because Terminal-Bench 2.1 tasks were fully disclosed, though it did not provide direct evidence that Google or Meta bought similar reinforcement learning environment data. Wang called the argument “very stupid,” saying Terminal-Bench 2.1 was already saturated and no longer effective at separating top-tier models, while Terminal-Bench 4.0 was not yet saturated.

Meta Chief AI Officer Alexandr Wang rejected a SemiAnalysis claim that Gemini 3.8 Flash and Muse Spark 1.3 were obvious examples of “benchmaxxed” models, or systems optimized too aggressively for public evaluations.

SemiAnalysis highlighted the two models’ performance on two versions of Terminal-Bench. On Terminal-Bench 2.1, Gemini 3.8 Flash scored 89.4% and Muse Spark 1.3 scored 88.8%, putting them close to GPT-6 Astra and Claude Fable 5.1. On the recently released Terminal-Bench 4.0, their scores fell to 19.1% and 33.3%. Astra and Fable 5.1, by comparison, posted 57.7% and 55.8% on the newer version.

The dispute centers on whether public tests are being targeted

SemiAnalysis said the issue may lie in how easy public benchmarks have become to optimize against. All tasks in Terminal-Bench 2.1 were publicly available. Even if model developers did not train directly on the test questions, the firm said they could still buy reinforcement learning environment data that closely resembled those tasks. In that case, models could become better and better at this specific exam without showing the same level of generalization on new tasks.

SemiAnalysis did not provide direct evidence that Google or Meta purchased that kind of data.

Wang says Terminal-Bench 2.1 is saturated

Wang pushed back soon after and called the argument “very stupid.” He pointed to GPT-5.6 Sol as another example, saying that model also dropped from 88.8% on Terminal-Bench 2.1 to 37.3% on version 4.0.

His explanation was that Terminal-Bench 2.1 had already become saturated and could no longer distinguish effectively among top-tier models, while Terminal-Bench 4.0 remained far from saturation.

Terminal-Bench 4.0 changed tasks and resource limits

According to the report, Terminal-Bench 4.0 removed eight tasks that were saturated, had public solutions, or had other issues. It also modified 19 tasks and reset limits for time, CPU, and memory resources.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
7700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.