Inference Efficiency: HBM Shrinks, CUDA Challenged

Inference is getting cheaper from several directions at once. HBM bandwidth stays the same no matter how tall the memory stack is. So 12-hi HBM costs 26% more than 4-hi while giving only about 10% more speed. On large rack systems, most of the memory sits unused anyway, making the extra cost hard to justify. Major AI chip teams are now designing HBM4 around the 4-hi conclusion. At the same time, local AI inference improved 18x in accuracy per joule over 16 months — 5.9x from hardware and 3x from model efficiency improvements. The gap between on-device and cloud AI is closing faster than benchmarks show.
On hardware competition: Japanese neocloud ai& launched on Tenstorrent chips in one of the largest such deployments outside the US. Its CEO says NVIDIA's CUDA ecosystem is no longer a moat. Positron AI raised $875M for Asimov, a chip using cheap LPDDR5X memory instead of HBM. It targets over 90% memory bandwidth use during token generation versus under 30% for typical GPU inference. Asimov has not yet taped out and its current product Atlas lacks independent performance verification.



