Azalia Mirhoseini
Stanford — AI agents / multi-agent systems research
Stanford
on X

Local AI models gained 18x accuracy-per-joule in 16 months: 5.9x from hardware, 3x from model efficiency.

September 11, 2026
brightray analysis
Summary

Local AI inference has improved 18x in accuracy per joule over 16 months, driven by a 5.9x gain from hardware advances and a 3x gain from model efficiency improvements. This quantifies the rapid convergence of on-device and cloud AI capabilities, highlighting that the gap is closing faster than most benchmarks have captured.

Why it matters
  • 18x accuracy-per-joule gain in 16 months (5.9x hardware + 3x model) quantifies how fast local inference is catching cloud-level capability.
  • The hardware and model gains are roughly multiplicative, suggesting both tracks compound independently — meaning the trend may continue.
  • The hybrid local/cloud inference framing signals a strategic shift: workloads once assumed to require cloud compute are now viable on-device.
View original on x.com

Community notes

No notes yet — be the first.


See every signal in the Feed