SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests

SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests

N
News Editor
2026-09-12 08:20:11
SemiAnalysis ignited a fresh dispute over AI model benchmarking on Sept. 8, accusing Google’s Gemini 3.8 Flash and Meta’s Muse Spark 1.3 of showing the clearest signs of leaderboard gaming. The claim centered on a sharp performance gap between Terminal-Bench 2.1, where both models ranked near the top, and Terminal-Bench 4.0, released on Aug. 29 with revised tasks, stricter scoring rules and anti-cheating checks. Gemini 3.8 Flash fell from 89.4 on Terminal-Bench 2.1 to 19.1 on version 4.0, while SemiAnalysis charted Muse Spark 1.3 at 33.3 on the newer benchmark. Meta Chief AI Officer Alexandr Wang pushed back almost immediately, calling the criticism “a stupid argument” and saying large score drops also affected other models, including GPT-5.6 Sol. Beyond the score changes, SemiAnalysis argued that the bigger issue is an emerging market for benchmark-adjacent training data and reinforcement learning environments, where labs can buy tasks designed to closely resemble public tests without using the original questions directly. The firm also singled out Datacurve, saying the company both operates the DeepSWE benchmark and sells expert coding data and RL environments to frontier labs, raising questions about conflicts of interest in public model evaluations.

SemiAnalysis on Sept. 8 published five posts on X accusing Google and Meta of fielding models with what it called the clearest signs of leaderboard gaming, naming Gemini 3.8 Flash and Muse Spark 1.3.

SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests 2

The argument rested on score gaps between two versions of Terminal-Bench. In Terminal-Bench 2.1, Gemini 3.8 Flash scored 89.4 and ranked No. 2 out of 182 models, ahead of GPT-6 Astra at 88.4. On Terminal-Bench 4.0, which went live on Aug. 29, the same model scored 19.1 and fell to No. 12 among 14 tested systems. GPT-6 Astra also dropped, but only from 88.4 to 57.7.

Meta Chief AI Officer Alexandr Wang responded in the comments 27 minutes later, writing, “This is a stupid argument.” He said GPT-5.6 Sol fell from 88.8 on Terminal-Bench 2.1 to 37.3 on version 4.0, a larger decline, yet no one was claiming Sol had gamed the benchmark.

Wang also said Meta had never claimed Muse Spark 1.3 was as strong as Astra or Fable 5.1. The company’s position, he said, was that Muse Spark 1.3 offered better cost performance.

SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests 3

Terminal-Bench 4.0 reshuffled the rankings

Terminal-Bench is designed to test how well an agent handles real work. Models are given a terminal and a vague objective, then must plan steps, call tools, write scripts and fix errors on their own.

On the older 2.1 benchmark, which had been widely used across the industry for months, Gemini 3.8 Flash scored 89.4 for second place. Muse Spark 1.3 scored 88.8 for fourth, followed by GPT-6 Astra at 88.4 and Claude Fable 5.1 at 85.02.

Terminal-Bench 4.0 changed both the test set and the scoring setup. The new version included 66 tasks, removed 8 older tasks that had already been “overfit,” heavily revised 19 others and added an 8-hour timeout across the board.

Its anti-cheating design also became stricter. Organizers introduced 35 scoring criteria and added an “adversarial cheating trial” in which agents were deliberately allowed to probe for loopholes in the reward system. Any task that proved vulnerable was discarded.

SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests 4

Once the new leaderboard came out, the order shifted sharply. Claude Mythos 5.1 (max) posted 60.9. GPT-6 Astra and Claude Fable 5.1 followed at 57.7 and 55.8, with Claude Opus 5 at 52.3. The previous-generation GPT-5.6 Sol and Terra scored 37.3 and 23.6. Gemini 3.8 Flash, by contrast, dropped to 19.1, while Muse Spark 1.3 was shown at 33.3 in SemiAnalysis’ chart.

That left Gemini 3.8 Flash in a very different position. On the old leaderboard it outscored GPT-6 Astra; on the new one, its score was less than one-third of Astra’s and also below GPT-5.6 Terra.

SemiAnalysis says buying benchmark-like data can mimic cheating

The second SemiAnalysis post pushed past leaderboard math and focused on the market for training data built around public evaluations.

SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests 5

Its view was that Terminal-Bench 2.1 tasks were fully public. Google and Meta, it argued, would not be reckless enough to train directly on the original questions, but they could buy data designed to closely match the benchmark’s task patterns. That would not be identical to leaking test answers into training, yet the practical outcome could look very similar.

SemiAnalysis described this as a more advanced form of benchmark gaming in 2026. Earlier contamination was cruder: mixing original benchmark items into pretraining corpora so models could memorize them. The newer approach, it said, is to buy “mock exams” that resemble public tests, usually in the form of closed task systems where models can trial actions and receive rewards.

The firm cited several prior examples. HumanEval had nearly 40% of its samples contaminated. GSM8K lost 13 points after decontamination. SWE-bench saw its score cut roughly in half when retested on a private codebase.

It also laid out current pricing. A single training task can sell for $200 to $2,000. Complex software engineering tasks can reach $20,000. Cloning a website into a UI training environment costs about $20,000. Recreating a product at Slack’s scale starts at $300,000. If a client wants exclusive rights, the price can rise another 4x to 5x.

SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests 6

Citing Epoch AI research, the article said Anthropic had discussed spending $1 billion a year in this area. Quarterly contracts often run into six or seven figures, and one researcher put the average at roughly $300,000 to $500,000 per quarter.

On the supply side, SemiAnalysis said at least 35 companies are in the business, most of them startups with fewer than 20 employees. It also noted that Scale AI had revenue above 1.4 billion before Meta took a stake in the company, while Surge’s ARR was approaching 1 billion.

Datacurve was singled out over benchmark and data roles

SemiAnalysis then named Datacurve directly.

SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests 7

On DeepSWE 1.1, a benchmark focused on long-horizon coding, Muse Spark 1.3 (max) ranked first with 75.4. GPT-6 Astra and Claude Opus 5 followed, while Gemini 3.8 Flash scored 73.8 for fourth place.

The issue, according to SemiAnalysis, is that Datacurve both runs the DeepSWE benchmark and sells “expert coding data and reinforcement learning environments” to frontier labs. In other words, the same company operates the test and also supplies products used in model training.

That overlap, the firm argued, puts conflict-of-interest concerns front and center in public model evaluation.

What happens when public benchmarks lose credibility

SemiAnalysis ended by widening the critique beyond a single leaderboard.

SemiAnalysis alleges Google and Meta models were benchmark-gamed on public AI tests 8

Its conclusion was that this is the eventual fate of all strong public benchmarks, and Terminal-Bench 4.0 is not exempt. The benchmark still looks reliable now, it said, only because it has been public for about two weeks. As long as tasks remain open, frontier labs will keep optimizing against them.

The proposed answer was more high-quality private benchmarks. But the article also warned about the cost of that shift. If evaluation moves fully behind closed doors, public leaderboards could sink into little more than marketing copy, while the most informative results circulate only between labs and private evaluators.

That may suit large labs, because no one can prepare in advance for a closed-book test. For developers, though, it would mean losing a common measuring stick for comparing models.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
2700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.