DeepSWE Benchmark Sees Two Leaders in One Night: Gemini Topped, Then Meta Spark 1.3 Hits 75.4%

DeepSWE Benchmark Sees Two Leaders in One Night: Gemini Topped, Then Meta Spark 1.3 Hits 75.4%

N
News Editor
2026-09-03 09:11:03
Google's Gemini 3.8 Flash scored 73.7% on DeepSWE v1.1, but Meta's Muse Spark 1.3 quickly surpassed it with 75.4%. The test evaluates AI agents handling long-cycle software engineering tasks across 113 real codebases. Meta's new version improved by 20.4 percentage points over its predecessor, though it hasn't yet entered the official leaderboard.

DeepSWE got two new leaders in the space of a few hours. Google’s Gemini 3.8 Flash hit 73.7% on DeepSWE v1.1, and the official board showed 74%±1%. That put it ahead of Claude Opus 5 and GPT-5.6 Sol. Then Meta came in later with Muse Spark 1.3. It scored 75.4% on the same test, lifting the public record by 1.7 percentage points.

DeepSWE asks AI agents to handle long-run software engineering work on their own across 113 real-world codebases. Google and Meta both used the mini-swe-agent framework. Google ran the high inference tier. Meta used the max tier. Meta’s earlier Spark 1.2 was at 55.0%. The new version jumped 20.4 percentage points. Fast move.

But Meta’s 75.4% has not been added to the official DeepSWE leaderboard yet. On Artificial Analysis’s Coding Agent Index, the limited preview Spark 1.3 max scored 68, second only to Claude Opus 5 xhigh. The public xhigh version scored 64.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.