Development · September 14, 2026

Cheap Models Now Match Frontier Performance

Three new model releases show a clear cost-performance shift. DeepSeek V4.1-Flash uses a new encoder-decoder design. It uses 8B active parameters for input processing and 16B for output. That cuts KV cache size by 8× and brings per-task cost to $0.27 versus $0.67 for V4 Pro. It still benchmarks above V4 Pro and supports 1M-token context under an MIT license. SWE-2 from Cognition is post-trained on Moonshot AI's 2.8T-parameter Kimi K3 using reinforcement learning. It matches Fable 5.1 on FrontierCode 1.1 at 64% lower cost.

Qwen3.8-Flash-Next and zAI GLM-5.3-Flash also benchmark above Claude Sonnet 5 and Opus 4.7 and can run locally. Together, these releases mean that for most production agent workloads, paying full frontier prices is no longer justified.

The 3 sources
Liked this?
Get the daily brief — the signal, distilled, each morning.
You’re in. Check your inbox.
Check spam and mark it not-spam so tomorrow’s lands too.