DeepSWE got two new leaders in the space of a few hours. Google’s Gemini 3.8 Flash hit 73.7% on DeepSWE v1.1, and the official board showed 74%±1%. That put it ahead of Claude Opus 5 and GPT-5.6 Sol. Then Meta came in later with Muse Spark 1.3. It scored 75.4% on the same test, lifting the public record by 1.7 percentage points.
DeepSWE asks AI agents to handle long-run software engineering work on their own across 113 real-world codebases. Google and Meta both used the mini-swe-agent framework. Google ran the high inference tier. Meta used the max tier. Meta’s earlier Spark 1.2 was at 55.0%. The new version jumped 20.4 percentage points. Fast move.
But Meta’s 75.4% has not been added to the official DeepSWE leaderboard yet. On Artificial Analysis’s Coding Agent Index, the limited preview Spark 1.3 max scored 68, second only to Claude Opus 5 xhigh. The public xhigh version scored 64.

