Cognition SWE-2: High-Value Coding Agent Model
Released by Cognition on 10 September 2026 and described as its most advanced coding model yet: it achieves 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper, pushing the Pareto frontier of capability and inference cost; SWE-2 is post-trained from the 2.8-trillion-parameter Kimi K3 and Cognition says it scaled RL to the multi-trillion-parameter regime for the first time, training all reasoning-effort levels in a single RL run, available from launch day in Devin Desktop and CLI and rolling out on Devin Web and Fusion
Tool Overview
Features, steps and FAQ below
Features
- ✓ Pareto frontier: Cognition says SWE-2 achieves 50.0% on FrontierCode 1.1 Main, within one point of the leading Fable 5.1 while being 64% cheaper, and that on FrontierCode 1.1 Main and DeepSWE 1.1 it beats SWE-1.7 and Grok 4.6 on both score and cost and matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price
- ✓ All reasoning effort levels in a single RL run: Cognition says SWE-2 is post-trained from the 2.8T-parameter Kimi K3, scaling RL to the multi-trillion-parameter regime for the first time, with the key addition being an RL algorithm that trains all reasoning-effort levels in one run and advances the whole cost-performance frontier
- ✓ Benchmark results: Cognition reports SWE-2 at 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1 and 27.3% on Terminal-Bench 4, versus Kimi K3's 44.2%, 68.5%, 88.3% and 21.5%, saying its RL adds 5-6 points on many benchmarks
- ✓ Smarter and cheaper: Cognition says SWE-2 medium scores higher than SWE-1.7 on FrontierCode 1.1 Main while taking 58% fewer turns and costing 81% less on average, with the efficiency gains coming from focused exploration: SWE-2 medium makes its first real edit after a median of 18 steps versus 48 for SWE-1.7
- ✓ Engineering behavior: Cognition says SWE-2 is better at writing end-to-end tests, more resourceful within the user's boundaries when the obvious path is blocked, and re-derives conclusions rather than re-asserting them when challenged, running artifacts to gather evidence; medium steps into action faster for simple tasks while high and max plan and verify more on complex ones
How to Use
- Use SWE-2 directly in Devin Desktop or the Devin CLI, where Cognition says it is available starting today
- Pick a reasoning effort level for the task: medium is faster and cheaper for simple and intermediate tasks, while high and max plan and verify more on complex ones
- Watch for the rollout on Devin Web and Fusion, which Cognition says SWE-2 is rolling out on
- Evaluate cost by completed-task cost rather than a single token rate, since SWE-2 is delivered through Devin products with no standalone API price list
FAQ
What is SWE-2?
Cognition says SWE-2 is its most advanced coding model yet, pushing the Pareto frontier of capability and inference cost and achieving 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper.
How was it trained?
Cognition says SWE-2 is post-trained from the 2.8T-parameter Kimi K3, scaling RL to the multi-trillion-parameter regime for the first time, with the key addition being an RL algorithm that trains all reasoning-effort levels in a single run.
What are its benchmark scores?
Cognition reports SWE-2 at 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1 and 27.3% on Terminal-Bench 4, adding 5-6 points over Kimi K3 on many benchmarks.
Is it cheaper than the previous generation?
Cognition says that on FrontierCode 1.1 Main, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average, making its first real edit after a median of 18 steps versus 48 for SWE-1.7.
How do I use SWE-2?
Cognition says SWE-2 is available starting today in Devin Desktop and CLI and is rolling out on Devin Web and Fusion; it is delivered through Devin products with no published standalone API price list, so evaluate it by completed-task cost.