
Cognition's SWE-2 matches Fable 5.1 on FrontierCode at 64% lower cost, post-trained via RL from Kimi K3's 2.8T parameters in a single run covering all effort levels.
September 12, 2026
brightray analysis
Summary
Cognition released SWE-2, a coding model post-trained with reinforcement learning on top of Moonshot AI's 2.8T-parameter Kimi K3. It scores 50.0% on FrontierCode 1.1 Main — within 1 point of Fable 5.1 — while costing 64% less, and is the first Cognition model to offer selectable reasoning-effort levels, all trained in a single RL run using slope-matched Pareto cost penalties. SWE-2 has no open weights or standalone API and runs only inside Devin.
Why it matters
- SWE-2 scores 50.0% on FrontierCode 1.1 Main vs Fable 5.1's 50.9%, at 64% lower cost — a new cost-efficiency frontier for frontier-class coding models.
- RL adds 5-6 points on top of the K3 base, showing substantial headroom even at the 2.8T-parameter scale.
- All three effort levels trained in one RL run via slope-matched Pareto cost penalties (R = S − λC) — a novel training-efficiency result.
- SWE-2 medium cuts turns 58% and cost 81% vs SWE-1.7 on the same benchmark, with median first-edit step dropping from 48 to 18.
- Clear weakness: Terminal-Bench 4 at 27.3%, roughly 30 points behind Fable 5.1 (55.8%) and GPT-6 Astra (57.9%), suggesting multi-step terminal reasoning remains a gap.
- No open weights, no API — Devin-only distribution limits broad research or deployment use.
Community notes—
No notes yet — be the first.
See every signal in the Feed