Robots Atlas>ROBOTS ATLAS
Infrastructure

HBM

2013ActivePublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
Replacing a wide planar memory bus (GDDR/DDR) with a vertical 3D stack of DRAM dies interconnected by TSVs and mounted on an interposer right next to the processor, delivering very high bandwidth (hundreds of GB/s to TB/s per stack) at lower energy per bit.
Category
Infrastructure
Abstraction level
Building block
Operation level
TrainingInferenceDeployment
Use cases
AI training accelerators (GPUs / TPUs)Large language model (LLM) inference and servingHigh-performance computing (HPC) and supercomputersData-center GPUsHigh-end graphics cardsHigh-bandwidth networking hardware

How it works

In HBM, 4 to 16 DRAM dies are stacked on top of each other and vertically connected by thousands of TSVs and microbumps, forming a very wide interface (1024-bit per stack, and 2048-bit from HBM4) split into many independent channels (e.g. 16 channels of 64 bits in HBM3). At the bottom of the stack sits a base/logic die with buffers and test logic that talks to the processor memory controller. The whole stack is placed next to the GPU/accelerator on a silicon interposer (2.5D packaging) that routes dense signal traces over a short distance. Because the bus is extremely wide, HBM reaches high aggregate bandwidth at a relatively low per-pin clock, lowering the energy per transferred bit compared with GDDR.

Problem solved

Memory bandwidth has become the main bottleneck of AI and HPC accelerators (the "memory wall"): traditional GDDR/DDR require wide on-PCB buses, consume significant power and cannot keep up with growing GPU compute. HBM shortens the memory-to-processor distance, multiplies the number of I/O lines and runs at lower per-pin clocks, delivering far higher bandwidth per watt and unlocking compute that would otherwise be starved by memory access.

Components

3D-stacked DRAM diesData storage

4 to 16 DRAM memory layers stacked vertically on top of each other, forming a single memory stack.

Through-Silicon Vias (TSV)Vertical interconnect between layers

Vertical electrical connections passing through the silicon die that link the stacked DRAM layers and provide a wide, short interface.

Base / logic dieStack-to-memory-controller interface

A die at the bottom of the stack containing buffer circuitry and test logic; it mediates between the DRAM layers and the processor memory controller.

MicrobumpsInter-die connections

Fine solder connections joining adjacent layers of the stack and the stack to the base die.

Silicon interposer (2.5D)HBM-stack-to-processor connection

A silicon substrate routing dense signal traces between the HBM stack and the GPU/CPU/accelerator over a very short distance (2.5D packaging).

Implementation

Implementation pitfalls
Advanced packaging and manufacturing yieldHigh

TSV stacks and interposer assembly (e.g. TSMC CoWoS) are expensive, and packaging throughput is often a supply bottleneck for AI accelerators.

Fix:Reserve packaging capacity ahead of time and diversify HBM suppliers (SK hynix, Samsung, Micron).
Limited, non-upgradeable capacityMedium

HBM is integrated on the processor package; capacity is fixed at manufacturing, smaller than DDR modules and cannot be expanded.

Fix:Choose the right stack-height variant at purchase time; combine with system memory (e.g. CXL/DDR) for larger capacity.
Stack thermal managementMedium

Dense, tall DRAM stacks (12–16 layers) impede heat dissipation and may force clock throttling.

Fix:Advanced package cooling and careful thermal design of the accelerator.

Evolution

Original paper · 2013 · JEDEC
High Bandwidth Memory (HBM) DRAM — JESD235
2013
First HBM standard published (JESD235)
Inflection point

JEDEC defines HBM: 1024-bit interface per stack, up to 4 DRAM layers, ~128 GB/s per stack.

2015
First commercial use (AMD Radeon R9 Fury X)

The AMD Fiji-based GPU is the first product to use first-generation HBM.

2016
HBM2 (JESD235A)

Up to 8 layers, 8 GB per stack and up to 256 GB/s; used in NVIDIA Tesla P100 and V100 among others.

2019
HBM2E update

Higher clocks (up to ~3.6 Gbps/pin), ~460 GB/s per stack and up to 24 GB (12 layers).

2022
HBM3 (JESD238)
Inflection point

16 channels of 64 bits, up to 6.4 Gbps/pin (~819 GB/s per stack); powers NVIDIA H100.

2024
HBM3E — volume production
Inflection point

SK hynix and Micron: >9.2 Gbps/pin and >1.2 TB/s per stack, 24–36 GB; powers NVIDIA H200 and Blackwell.

2025
HBM4 (JESD270-4)
Inflection point

Doubled interface to 2048 bits per stack, up to 64 GB (16 layers); foundation of the NVIDIA Rubin platform.

Hyperparameters (configurable axes)

Stack height (number of DRAM dies)Critical

Number of DRAM layers in the stack — affects capacity and partly bandwidth.

4-high (HBM)
8-high (HBM2)
12-high (HBM2E/HBM3)
16-high (HBM4)
Interface width per stackCritical

Width of a single stack bus in bits.

1024-bit (HBM – HBM3E)
2048-bit (HBM4)
Capacity per stackHigh

Maximum capacity of a single HBM stack.

4 GB (HBM)
8 GB (HBM2)
24 GB (HBM2E / HBM3)
36 GB (HBM3E)
64 GB (HBM4)
Per-pin data rateHigh

Transfer speed per single I/O pin.

~1 Gbps (HBM)
~2 Gbps (HBM2)
~3.6 Gbps (HBM2E)
~6.4 Gbps (HBM3)
>9.2 Gbps (HBM3E)
Number of channelsMedium

Division of the wide stack bus into independent channels.

8 × 128-bit (HBM2E)
16 × 64-bit (HBM3)

Hardware requirements

Primary

Modern AI GPUs (NVIDIA H100/H200/Blackwell) are memory-bandwidth bound; HBM supplies the TB/s that tensor cores need.

Good fit

Google TPU accelerators use HBM as main memory to feed large matrix-multiply units.