Meta Chief AI Officer Alexandr Wang rejected a SemiAnalysis claim that Gemini 3.8 Flash and Muse Spark 1.3 were obvious examples of “benchmaxxed” models, or systems optimized too aggressively for public evaluations.
SemiAnalysis highlighted the two models’ performance on two versions of Terminal-Bench. On Terminal-Bench 2.1, Gemini 3.8 Flash scored 89.4% and Muse Spark 1.3 scored 88.8%, putting them close to GPT-6 Astra and Claude Fable 5.1. On the recently released Terminal-Bench 4.0, their scores fell to 19.1% and 33.3%. Astra and Fable 5.1, by comparison, posted 57.7% and 55.8% on the newer version.
The dispute centers on whether public tests are being targeted
SemiAnalysis said the issue may lie in how easy public benchmarks have become to optimize against. All tasks in Terminal-Bench 2.1 were publicly available. Even if model developers did not train directly on the test questions, the firm said they could still buy reinforcement learning environment data that closely resembled those tasks. In that case, models could become better and better at this specific exam without showing the same level of generalization on new tasks.
SemiAnalysis did not provide direct evidence that Google or Meta purchased that kind of data.
Wang says Terminal-Bench 2.1 is saturated
Wang pushed back soon after and called the argument “very stupid.” He pointed to GPT-5.6 Sol as another example, saying that model also dropped from 88.8% on Terminal-Bench 2.1 to 37.3% on version 4.0.
His explanation was that Terminal-Bench 2.1 had already become saturated and could no longer distinguish effectively among top-tier models, while Terminal-Bench 4.0 remained far from saturation.
Terminal-Bench 4.0 changed tasks and resource limits
According to the report, Terminal-Bench 4.0 removed eight tasks that were saturated, had public solutions, or had other issues. It also modified 19 tasks and reset limits for time, CPU, and memory resources.

