Cognition released SWE-2 on 10 September 2026, and one number travelled: 50.0% on FrontierCode 1.1 Main, within a single point of Anthropic's Fable 5.1, at 64% lower cost. That number is real, and it comes from Cognition's own benchmark table.
Three rows below it, in the same table, on the same page, is a number that did not travel. On Terminal-Bench 4, SWE-2 scores 27.3%. Fable 5.1 scores 55.8%. GPT-6 Astra scores 57.9%.
Cognition did not bury this. It published both figures in one four-row table and let readers draw the line. The coverage drew a different one. What follows reads the whole table, checks which benchmark versions are actually current, and works out what a builder can do with SWE-2 today, which is narrower than the launch framing suggests.
What Cognition shipped
SWE-2 is Cognition's flagship coding model, the successor to SWE-1.7 from July. It is post-trained from Kimi K3, Moonshot AI's 2.8T-parameter open-weights model, using reinforcement learning on top of the SWE-1.7 training infrastructure.
The technical headline is the training method rather than the model. Cognition trains all three reasoning-effort levels, medium, high and max, inside a single RL run, using a cost-penalised reward of the form R = S minus lambda times C, where the penalty for each effort level is tuned to the local slope of the base model's cost-performance frontier. The company argues from first principles that only a linear cost penalty gives the same answer whether you apply it before or after averaging, and proves it in an appendix.
Availability is the part that matters this week, and it is stated plainly: SWE-2 is available "starting today in Devin Desktop and CLI", with Devin Web and Fusion still rolling out. There is no separate model endpoint.

The full table, including the row nobody quoted
Every figure below is Cognition's, reproduced in full rather than filtered. Higher is better in all four rows.
| Benchmark | SWE-2 | Kimi K3 | Grok 4.6 | Fable 5.1 | GPT-5.6 Sol | GPT-6 Astra | SWE-1.7 |
|---|---|---|---|---|---|---|---|
| FrontierCode 1.1 Main | 50.0% | 44.2% | 48.0% | 50.9% | 47.5% | 53.3% | 42.0% |
| DeepSWE 1.1 | 73.0% | 68.5% | 67.5% | 67.4% | 72.7% | 74.1% | 37.7% |
| Terminal-Bench 2.1 | 92.8% | 88.3% | 88.4% | 91.4% | 88.8% | 89.9% | 81.5% |
| Terminal-Bench 4 | 27.3% | 21.5% | 20.3% | 55.8% | 37.3% | 57.9% | 7.6% |
One methodology note belongs with the table, and it is in Cognition's Appendix A rather than the summary. Where a public result existed, Cognition used it. Where one did not, Cognition ran the model itself, using the harness each model was built for: Claude Code for Anthropic models, Codex for OpenAI models, Grok Build for xAI models, and Devin CLI for open-weight models. That last clause means Kimi K3, the closest comparison in the table and SWE-2's own base, was scored inside Cognition's product. Every model is reported at its best score across reasoning-effort settings.
The Terminal-Bench split, and which version is live
On Terminal-Bench 2.1, SWE-2 does not merely compete. It tops the table at 92.8%, ahead of Fable 5.1 at 91.4% and GPT-6 Astra at 89.9%. On Terminal-Bench 4, it lands at less than half of Fable 5.1.
Same benchmark family, two versions, opposite verdicts. So which version is the one people are running now? The Terminal-Bench leaderboard, hosted by Stanford, Harbor and the Laude Institute, currently ranks models on Terminal-Bench 4.0. On GitHub, the Terminal-Bench 4 repository was last pushed on 3 September 2026, while the separate 2.1 repository was last pushed on 26 August. Version 4 is the live benchmark; 2.1 is the superseded one.
That inverts the headline framing. The Terminal-Bench row where SWE-2 leads the field is the retired version. The row where it trails Fable 5.1 by 28.5 points is the current one.
The fair counterpoint is that every model falls on Terminal-Bench 4, so the benchmark is simply harder. True, and worth stating. But the falls are not proportional. Fable 5.1 drops 35.6 points between the two versions and GPT-6 Astra drops 32.0. SWE-2 drops 65.5. Kimi K3 drops 66.8.

The arithmetic that explains the whole table
Line SWE-2 up against Kimi K3, its own base model, on all four benchmarks and the gap barely moves:
- FrontierCode 1.1 Main: 44.2% to 50.0%, a gain of 5.8 points
- DeepSWE 1.1: 68.5% to 73.0%, a gain of 4.5 points
- Terminal-Bench 2.1: 88.3% to 92.8%, a gain of 4.5 points
- Terminal-Bench 4: 21.5% to 27.3%, a gain of 5.8 points
Cognition states this itself, writing that its RL adds 5 to 6 points on many benchmarks. What the company does not spell out is the consequence: post-training moved the level, not the shape. Wherever Kimi K3 is strong, SWE-2 is strong plus roughly five. Wherever Kimi K3 is weak, SWE-2 is weak plus roughly five. Terminal-Bench 4 is not a place SWE-2 uniquely struggles. It is a place its base model struggles, and five points does not close a 34-point hole.
For anyone choosing a model, that is a usable heuristic rather than a criticism. If you want to predict how SWE-2 will handle a task category Cognition did not benchmark, look up how Kimi K3 handles it and add five points. That is a more reliable forecast than the frontier comparison in the headline.
You are buying Kimi K3 plus five points, and Cognition knows the question that raises
The base model is Chinese open weights. Kimi K3 sits on Hugging Face with roughly 2.3 million downloads and more than 11,000 likes, under a custom licence Moonshot names "kimi-k3" rather than a standard open-source licence, and it is described in the Kimi K3 paper from July 2026.
Cognition met the obvious objection with data rather than assurance, reusing the framework from its earlier work on measuring the trustworthiness of open-source-derived models. It re-ran a propaganda and censorship evaluation built on 145 questions about politically sensitive topics in China, submitted in English, Simplified Chinese and Traditional Chinese. SWE-2 passed 98.0% of attempts overall: 99.8% in English, 95.2% in Simplified Chinese, 99.1% in Traditional Chinese.
A second evaluation tested whether customer identity or request language changes a model's willingness to implement vulnerable or abusive functionality, using Western, Pakistani, Chinese, Tibetan and Falun Gong customer framings with some requests written in Urdu or Chinese. Cognition reports that no framing condition produced a statistically significant change in vulnerability for any of the six models tested, which included Kimi K3, GLM 5.3, GPT 5.6, Fable 5.1 and Opus 5 alongside SWE-2.
Running that suite against your own model, publishing the per-language breakdown, and including your own base model in the comparison is more disclosure than most launches carry.

What you cannot buy
Cognition published no standalone API for SWE-2, no price for it, and no weights. The only way to run the model is a Devin subscription, and at launch only two of the four Devin surfaces have it.
This makes the cost claim awkward in an interesting way. Cognition's footnote gives precise per-task dollar figures for its competitors: Fable 5.1 Max scores 50.3% at $12.83 per task on FrontierCode 1.1 Main against Fable 5.1 Medium's 50.9% at $3.28, and Fable 5 Max scores 69.7% at $21.63 per task on DeepSWE 1.1 against Fable 5 xhigh's 69.9% at $13.41. There is no equivalent published dollar figure for SWE-2 itself. The 64% saving is a relative claim measured against absolute prices you can look up for the other side only.
The practical shape of the market, then, is that Fable 5.1 and GPT-6 Astra are models you can call, Kimi K3 is a model you can download, and SWE-2 is neither. It is a property of a product. Cognition's Devin documentation is where the surfaces and limits are described, and it is worth reading before treating SWE-2 as a drop-in swap for an API model in an existing pipeline.
The efficiency change is the part you will actually feel
Underneath the benchmark argument sits the improvement most likely to change a working day, and it is measured against SWE-1.7 rather than against the frontier.
On FrontierCode 1.1 Main, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average. Mean steps per run fall from 127 for SWE-1.7 to 53 for SWE-2 medium, with SWE-2 high at 80 and SWE-2 max at 98. The model reaches its first real edit after a median of 18 steps, against 48 for SWE-1.7.
Cognition attributes this to focused exploration rather than shorter thinking: user feedback said SWE-1.7 over-explored and overthought simple tasks, and the fix was judgment about which parts of a codebase matter. Whether that holds outside the 100-task FrontierCode set is not something this article tested. But the direction is the right one for anyone who has watched an agent read forty files before touching one.
Cognition also reports behavioural changes it observed internally rather than measured: better end-to-end test writing, more willingness to find another route when a path is blocked, and a habit of re-deriving conclusions when challenged instead of agreeing. Those are qualitative claims from the vendor, presented as such.

How to decide this week
If you already pay for Devin, this is a straightforward update: SWE-2 is in Desktop and CLI now, and medium is the level Cognition's own step counts point at for routine work, with high and max held back for tasks where the plan matters more than the turnaround.
If you are choosing a coding model and you are not on Devin, understand what the decision actually is. You cannot buy SWE-2, so the choice is whether to adopt Devin as a platform. The benchmark comparison in the announcement is a comparison between a product and two APIs.
If you self-host, the honest read of the arithmetic above is that Kimi K3 gets you within about five points on every published benchmark, and it is downloadable today under Moonshot's own licence terms, which are worth reading before commercial use.
And whatever you pick, take the version numbers seriously. A vendor table that reports both Terminal-Bench 2.1 and Terminal-Bench 4 is doing the right thing. A summary that quotes only the version where the vendor wins is not.
Frequently asked questions
Is SWE-2 available through an API?
No. Cognition has not published a standalone API, pricing, or weights for SWE-2. At launch it runs in Devin Desktop and Devin CLI, with Devin Web and Fusion described as rolling out.
What model is SWE-2 built on?
SWE-2 is post-trained from Kimi K3, Moonshot AI's 2.8T-parameter model, which Cognition notes had already undergone extensive reinforcement learning for agentic coding before Cognition's own RL was applied on top.
Why does SWE-2 score 92.8% on Terminal-Bench 2.1 but 27.3% on Terminal-Bench 4?
Terminal-Bench 4 is a harder and newer version of the benchmark, and every model in Cognition's table scores lower on it. The gap is that SWE-2 falls further than the frontier models do: 65.5 points between versions, against 35.6 for Fable 5.1 and 32.0 for GPT-6 Astra. Version 4.0 is the one currently ranked on the public Terminal-Bench leaderboard.
Is SWE-2 better than Fable 5.1?
It depends entirely on which row you read. SWE-2 leads on DeepSWE 1.1 (73.0% against 67.4%) and on Terminal-Bench 2.1 (92.8% against 91.4%). It trails narrowly on FrontierCode 1.1 Main (50.0% against 50.9%) and heavily on Terminal-Bench 4 (27.3% against 55.8%). Cognition's claim is cost-adjusted parity, not outright superiority.
How much does SWE-2 cost?
Cognition states that SWE-2 is 64% cheaper than Fable 5.1 at a comparable FrontierCode 1.1 Main score, but published no absolute price for SWE-2. Costs in the announcement assume list pricing including public discounts, and the only dollar figures given are per-task costs for competing models.
What are SWE-2's effort levels and which should I use?
Medium, high and max, all trained in one RL run rather than as separate models. Cognition reports medium stepping into action fastest, at a mean of 53 steps per run, with high at 80 and max at 98 and both holding an edge on complex tasks through more planning and verification.
Are SWE-2's weights open?
No. The base model, Kimi K3, is available on Hugging Face under a custom Moonshot licence, but SWE-2's post-trained weights have not been released.