Skip to main content

Cognition SWE-2 Reaches Devin With a 64% Cost Claim and One Benchmark Warning

3 min read

Cognition SWE-2 is reaching Devin with a company-reported 64% cost advantage in one comparison. The wider benchmark table tells a less simple story.

Cognition SWE-2 Reaches Devin With a 64% Cost Claim and One Benchmark Warning

Cognition has a sharp headline for SWE-2: nearly the same FrontierCode score as Fable 5.1 at 64% lower cost. Then you move one row down the benchmark table and find a much larger gap on Terminal-Bench 4. Both numbers matter.

Cognition introduced SWE-2 on September 10, 2026. The model is available in Devin Desktop and the Devin CLI. Cognition says it is rolling out to Devin Web and Fusion.

The model started from Kimi K3

Cognition says SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter model that had already received agentic coding reinforcement learning. Cognition then trained multiple reasoning-effort levels in one run, with a reward designed to improve the cost-performance frontier instead of optimizing only the maximum score.

That is a useful product goal. Most developers do not buy a benchmark percentage. They buy completed changes under a budget and deadline.

The comparison table needs more than one row

BenchmarkSWE-2Fable 5.1GPT-6 AstraSource status
FrontierCode 1.1 Main50.0%50.9%53.3%Cognition-reported
DeepSWE 1.173.0%67.4%74.1%Cognition-reported
Terminal-Bench 2.192.8%91.4%89.9%Cognition-reported
Terminal-Bench 427.3%55.8%57.9%Cognition-reported
These values come from Cognition’s launch post. They are company-reported results, not independent MustHave.ai tests.

SWE-2 looks competitive on FrontierCode, DeepSWE and Terminal-Bench 2.1 in Cognition’s table. The Terminal-Bench 4 row is the warning. It shows why “frontier” depends on the task mix and harness, not only the model name.

The gaps make that warning concrete. SWE-2 trails Fable 5.1 by only 0.9 percentage points on FrontierCode 1.1 Main, the row used for Cognition’s cost comparison. On Terminal-Bench 4, the gap is 28.5 points. Both differences are calculated from Cognition’s own table, and they describe different task distributions rather than a contradiction.

For a 100-task internal trial, a one-point score difference may amount to roughly one additional solved task, while a 28.5-point difference would be large enough to change routing policy. That is why a team should report results by task family before averaging them into one model score.

The 64% figure has a boundary

Cognition’s 64% cheaper claim compares SWE-2 with Fable 5.1 near the same FrontierCode score. It is not a universal discount across every task, effort setting or product plan. The company has not published a standalone SWE-2 API price in the launch post.

Cognition also reports that SWE-2 medium used 53 mean steps across the 100-task FrontierCode set, versus 127 for SWE-1.7, and made its first real edit after a median 18 steps instead of 48. Those are internal measurements with three runs per task.

Cost per accepted change is the metric to copy

A cheap run that fails review is expensive. A longer run can be worth paying for if it removes debugging and rework. Teams should measure the complete loop: model charge, sandbox time, retries, human review, CI and rollback.

  • Use tasks from your own repositories, not only public benchmark prompts.
  • Run each task more than once at every effort level you may buy.
  • Count accepted changes after tests and review, not agent self-reports.
  • Record total cost and elapsed time for failures as well as successes.
  • Keep the harness, tools and repository snapshot fixed during comparison.

The same discipline appears in our AI agent cost-control guide. For a broader product comparison, see Codex versus Claude Code with Astra and Fable.

What is available and what is missing

SWE-2 is a Devin model today, not a generally documented standalone API product. Cognition did not publish weights, a context-window specification, a model card with independent replication or a separate public price sheet in the materials reviewed.

That does not make the release unimportant. It changes the buying question from “Which model tops the chart?” to “Which Devin configuration completes my accepted task at the lowest total cost?”

My take: the weak row is the useful row

I like that Cognition published enough of the table to expose a weakness. The Terminal-Bench 4 result prevents the launch from collapsing into one neat marketing number.

If your workload resembles repository maintenance, shell-heavy operations or full software tasks, build a mixed evaluation set. Let SWE-2 win where it wins. Keep another model for the jobs it does not.

Read the primary material

Checked September 12, 2026. Every score, cost comparison and behavior claim is attributed to Cognition unless separately stated.

Leave a comment

Your email address will not be published. Required fields are marked *