← Back to AI Tools

Cognition SWE-2: High-Value Coding Agent Model

Released by Cognition on 10 September 2026 and described as its most advanced coding model yet: it achieves 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper, pushing the Pareto frontier of capability and inference cost; SWE-2 is post-trained from the 2.8-trillion-parameter Kimi K3 and Cognition says it scaled RL to the multi-trillion-parameter regime for the first time, training all reasoning-effort levels in a single RL run, available from launch day in Devin Desktop and CLI and rolling out on Devin Web and Fusion

Tool Overview

Features, steps and FAQ below

Features

  • ✓ Pareto frontier: Cognition says SWE-2 achieves 50.0% on FrontierCode 1.1 Main, within one point of the leading Fable 5.1 while being 64% cheaper, and that on FrontierCode 1.1 Main and DeepSWE 1.1 it beats SWE-1.7 and Grok 4.6 on both score and cost and matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price
  • ✓ All reasoning effort levels in a single RL run: Cognition says SWE-2 is post-trained from the 2.8T-parameter Kimi K3, scaling RL to the multi-trillion-parameter regime for the first time, with the key addition being an RL algorithm that trains all reasoning-effort levels in one run and advances the whole cost-performance frontier
  • ✓ Benchmark results: Cognition reports SWE-2 at 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1 and 27.3% on Terminal-Bench 4, versus Kimi K3's 44.2%, 68.5%, 88.3% and 21.5%, saying its RL adds 5-6 points on many benchmarks
  • ✓ Smarter and cheaper: Cognition says SWE-2 medium scores higher than SWE-1.7 on FrontierCode 1.1 Main while taking 58% fewer turns and costing 81% less on average, with the efficiency gains coming from focused exploration: SWE-2 medium makes its first real edit after a median of 18 steps versus 48 for SWE-1.7
  • ✓ Engineering behavior: Cognition says SWE-2 is better at writing end-to-end tests, more resourceful within the user's boundaries when the obvious path is blocked, and re-derives conclusions rather than re-asserting them when challenged, running artifacts to gather evidence; medium steps into action faster for simple tasks while high and max plan and verify more on complex ones

How to Use

  1. Use SWE-2 directly in Devin Desktop or the Devin CLI, where Cognition says it is available starting today
  2. Pick a reasoning effort level for the task: medium is faster and cheaper for simple and intermediate tasks, while high and max plan and verify more on complex ones
  3. Watch for the rollout on Devin Web and Fusion, which Cognition says SWE-2 is rolling out on
  4. Evaluate cost by completed-task cost rather than a single token rate, since SWE-2 is delivered through Devin products with no standalone API price list

FAQ

What is SWE-2?

Cognition says SWE-2 is its most advanced coding model yet, pushing the Pareto frontier of capability and inference cost and achieving 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper.

How was it trained?

Cognition says SWE-2 is post-trained from the 2.8T-parameter Kimi K3, scaling RL to the multi-trillion-parameter regime for the first time, with the key addition being an RL algorithm that trains all reasoning-effort levels in a single run.

What are its benchmark scores?

Cognition reports SWE-2 at 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1 and 27.3% on Terminal-Bench 4, adding 5-6 points over Kimi K3 on many benchmarks.

Is it cheaper than the previous generation?

Cognition says that on FrontierCode 1.1 Main, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average, making its first real edit after a median of 18 steps versus 48 for SWE-1.7.

How do I use SWE-2?

Cognition says SWE-2 is available starting today in Devin Desktop and CLI and is rolling out on Devin Web and Fusion; it is delivered through Devin products with no published standalone API price list, so evaluate it by completed-task cost.

Official Source

https://cognition.com/blog/swe-2