What is HBM and why does it matter?

AIAcademy · AIAcademy · 2026-05-16

Micron HBM memory

The binding constraint on AI accelerators in 2026 is not FLOPS. It is High-Bandwidth Memory — the stacks of DRAM chips that sit millimeters from the GPU die and feed it tokens. Double the FLOPS without doubling the memory bandwidth and you get very little: the compute sits idle, waiting for weights and KV cache to arrive.

What HBM actually is. A vertical stack of DRAM dies wired together with thousands of through-silicon vias, sitting on a silicon interposer next to the GPU. Where conventional GDDR memory connects over a few hundred bits at high frequency, HBM connects over a much wider interface next to the accelerator. Wider and closer beats faster and farther. The current frontier transition is from HBM3E into HBM4: NVIDIA's Rubin GPU lists 288 GB of HBM4 and up to 22 TB/s of bandwidth, while AMD's MI400 preview lists up to 432 GB of HBM4 and 19.6 TB/s. A frontier training run is fundamentally a memory-bandwidth problem dressed up as a compute problem.

Why the supply is a three-company story. SK hynix, Samsung Semiconductor, and Micron are the three public HBM4 supplier lanes learners should know. Their source pages talk about I/O width, stack height, bandwidth, power efficiency, and production readiness because accelerator designers qualify memory as part of the whole package. There is no simple late swap from HBM4 to commodity DDR or GDDR without changing the performance class of the system.

Why context windows make this worse. Nvidia's Rubin CPX announcement is the cleanest signal: a dedicated GPU class for million-token-context inference, because KV cache at long context is the dominant memory footprint, and the GPU you want for that workload is mostly memory and interconnect with comparatively modest compute. The industry has stopped pretending FLOPS is the headline number.