
4-hi HBM stacks offer identical bandwidth to 12-hi at ~33% of the GB cost, making them optimal for inference TCO as rack-scale systems already over-provision capacity.
Discovered September 13, 2026 · Published September 13, 2026
brightray analysis
Summary
HBM bandwidth is constant regardless of stack height since all 2048 I/Os are accessible at 4-hi; cost scales with GB, not bandwidth. On rack-scale systems like Rubin Ultra NVL576 (21TB aggregate HBM), frontier model weights consume under 10% of capacity, leaving most memory stranded. Analysis of Kimi K3 inference shows 12-hi HBM costs 26.3% more than 4-hi but delivers only ~10% more throughput, resulting in higher cost-per-token — making 4-hi the optimal inference SKU, a configuration major lab ASIC teams are now targeting for HBM4 onwards.
Why it matters
- 4-hi HBM delivers the same per-cube bandwidth as 12-hi (2048 I/Os accessible at 4-hi) but at ~33% of the GB cost, making it strictly better $/bandwidth for inference.
- On Rubin Ultra NVL576 (21TB aggregate HBM), Kimi K3 weights occupy under 10% of capacity — stranded capacity is now the norm, not the exception, in rack-scale deployments.
- Quantified TCO impact: 12-hi costs 26.3% more than 4-hi but yields only ~10% more throughput; 8-hi costs 12.1% more for ~8% gain — both stacks deliver higher cost-per-token than 4-hi.
- Nvidia's Rubin Ultra already reflects this shift, cutting HBM from 288GB to 192GB per GPU; lab ASIC programs are targeting HBM4 4-hi for inference from next generation onwards.
- KVCache offloading to second-tier DRAM further reduces HBM capacity requirements, extending the economic case for shorter stacks even in high-concurrency agentic workloads.
- The broader implication: HBM wafer scarcity makes maximizing tokens-per-HBM-wafer as important as tokens-per-Watt, and 4-hi is the direct lever for both.
Community notes—
No notes yet — be the first.
See every signal in the Feed