Development · September 14, 2026

Inference Efficiency: HBM Shrinks, CUDA Challenged

Inference is getting cheaper from several directions at once. HBM bandwidth stays the same no matter how tall the memory stack is. So 12-hi HBM costs 26% more than 4-hi while giving only about 10% more speed. On large rack systems, most of the memory sits unused anyway, making the extra cost hard to justify. Major AI chip teams are now designing HBM4 around the 4-hi conclusion. At the same time, local AI inference improved 18x in accuracy per joule over 16 months — 5.9x from hardware and 3x from model efficiency improvements. The gap between on-device and cloud AI is closing faster than benchmarks show.

On hardware competition: Japanese neocloud ai& launched on Tenstorrent chips in one of the largest such deployments outside the US. Its CEO says NVIDIA's CUDA ecosystem is no longer a moat. Positron AI raised $875M for Asimov, a chip using cheap LPDDR5X memory instead of HBM. It targets over 90% memory bandwidth use during token generation versus under 30% for typical GPU inference. Asimov has not yet taped out and its current product Atlas lacks independent performance verification.

The 4 sources
Liked this?
Get the daily brief — the signal, distilled, each morning.
You’re in. Check your inbox.
Check spam and mark it not-spam so tomorrow’s lands too.