Cognition rolls out SWE-2, narrowing the gap with Astra at roughly a quarter of the cost

Cognition rolls out SWE-2, narrowing the gap with Astra at roughly a quarter of the cost

N
News Editor
2026-09-11 00:00:44
Cognition, the company behind coding agent Devin, has introduced a new programming model called SWE-2. The model is built by continuing reinforcement learning on Moonshot AI’s 2.8 trillion-parameter Kimi K3, and Cognition said the extra training lifted results across several coding benchmarks by about 5 to 6 percentage points. On Cognition’s in-house FrontierCode 1.1 Main benchmark, SWE-2 scored 50.0%, ahead of GPT-5.6 Sol at 47.5% and Grok 4.6 at 48.0%. It trailed Fable 5.1 by 0.9 percentage points, with Fable posting 50.9%, while running at 64% lower cost. GPT-6 Astra scored 53.3%, leaving SWE-2 3.3 percentage points behind, with Cognition putting SWE-2’s cost at about one-quarter of Astra’s. The company also said SWE-2 is more efficient than its predecessor: at medium reasoning intensity, task-completion interaction rounds fell 58% and average cost dropped 81%. Results were weaker on the harder Terminal-Bench 4, where SWE-2 scored 27.3%, well below Astra’s 57.9% and Fable 5.1’s 55.8%. SWE-2 is now live in Devin Desktop and CLI, and is starting to roll out to Devin Web and Fusion.

Cognition, the parent company of coding agent Devin, has released a new programming model, SWE-2. The model continues reinforcement learning on Moonshot AI’s 2.8 trillion-parameter Kimi K3, and Cognition said the added training improved results across several coding benchmarks by about 5 to 6 percentage points.

FrontierCode 1.1 Main results

On Cognition’s own FrontierCode 1.1 Main benchmark, SWE-2 posted a score of 50.0%, beating GPT-5.6 Sol at 47.5% and Grok 4.6 at 48.0%.

SWE-2 was 0.9 percentage points behind Fable 5.1, which scored 50.9%, while coming in at 64% lower cost, according to the comparison provided by Cognition.

GPT-6 Astra scored 53.3% on the same benchmark. That left SWE-2 trailing by 3.3 percentage points, with cost at about one-quarter of Astra’s.

Fewer interaction rounds and lower cost

Cognition said SWE-2 takes fewer detours than the previous generation. At medium reasoning intensity, the number of interaction rounds needed to complete tasks fell 58%, while average cost dropped 81%.

Gap remains on harder benchmark

On the more difficult Terminal-Bench 4, SWE-2 scored 27.3%. That was well below Astra’s 57.9% and Fable 5.1’s 55.8%, showing it has not yet fully caught up with frontier models across the board.

Availability

SWE-2 is already available on Devin Desktop and CLI, and is starting to roll out to Devin Web and Fusion.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.