MarkTechPost
AI-research news outlet — own-channel reporting
MarkTechPost
on marktechpost.com

Cognition's SWE-2 matches Fable 5.1 on FrontierCode at 64% lower cost, post-trained via RL from Kimi K3's 2.8T parameters in a single run covering all effort levels.

September 12, 2026
brightray analysis
Summary

Cognition released SWE-2, a coding model post-trained with reinforcement learning on top of Moonshot AI's 2.8T-parameter Kimi K3. It scores 50.0% on FrontierCode 1.1 Main — within 1 point of Fable 5.1 — while costing 64% less, and is the first Cognition model to offer selectable reasoning-effort levels, all trained in a single RL run using slope-matched Pareto cost penalties. SWE-2 has no open weights or standalone API and runs only inside Devin.

Why it matters
  • SWE-2 scores 50.0% on FrontierCode 1.1 Main vs Fable 5.1's 50.9%, at 64% lower cost — a new cost-efficiency frontier for frontier-class coding models.
  • RL adds 5-6 points on top of the K3 base, showing substantial headroom even at the 2.8T-parameter scale.
  • All three effort levels trained in one RL run via slope-matched Pareto cost penalties (R = S − λC) — a novel training-efficiency result.
  • SWE-2 medium cuts turns 58% and cost 81% vs SWE-1.7 on the same benchmark, with median first-edit step dropping from 48 to 18.
  • Clear weakness: Terminal-Bench 4 at 27.3%, roughly 30 points behind Fable 5.1 (55.8%) and GPT-6 Astra (57.9%), suggesting multi-step terminal reasoning remains a gap.
  • No open weights, no API — Devin-only distribution limits broad research or deployment use.
View original on marktechpost.com

Community notes

No notes yet — be the first.


See every signal in the Feed