Thorsten Ball
Engineer (Amp); author 'How to Build an Agent'
Amp
on registerspill.thorstenball.com

Fable 5.1 and GPT-6 Astra mark a qualitative shift: multi-agent orchestration — agents spawning and blind-testing agents — now works reliably in real production workflows.

September 12, 2026
brightray analysis
Summary

The latest frontier models (Fable 5.1, GPT-6 Astra) represent a qualitative capability jump where the previously failing 5% of complex agentic tasks is now reliably completed. Multi-agent patterns — agents spawning sub-agents, blind black-box testing, cross-codebase deployment verification — are functioning in real production use. The Navier-Stokes proof used 10,000 agents and ~300B output tokens; ARC-AGI benchmark cost collapsed from ~$500K (o3) to ~$20 (Astra) for comparable or higher scores.

Why it matters
  • Latest models now reliably complete the final 5% of complex agentic tasks, enabling fully delegated end-to-end workflows where the model provides video proof of correctness.
  • Multi-agent orchestration is production-ready: a lead agent spawns sub-agents, withholds context to blind-test them, evaluates their outputs, and self-corrects AGENTS.md — all autonomously.
  • ARC-AGI benchmark cost collapsed from ~$500K (o3, 87.5%) to ~$20 (Astra, higher score), suggesting Navier-Stokes-scale compute may cost ~$50 in a few years.
  • OpenAI pausing $200/mo Pro plan signups and Anthropic's Head of Platform citing real CPU shortages signal acute compute scarcity despite demand.
  • Terence Tao warns AI's ability to rapidly flatten research problems may incentivize secrecy, reversing centuries of open-science tradition.
  • An Anthropic researcher's resignation over self-improving superintelligence risk drew 165M views, with a colleague publicly assigning >10% probability of AI killing all humans within a decade.
View original on registerspill.thorstenball.com

Community notes

No notes yet — be the first.


See every signal in the Feed