Robots Atlas>ROBOTS ATLAS
Inference

Memory-bound decoding

2009ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Capturing that autoregressive token generation is limited by memory bandwidth (not compute), because each decode step must read all model weights and the KV cache from memory while performing few operations on them — at low arithmetic intensity.
Category
Inference
Abstraction level
Pattern
Operation level
InferenceServing
Use cases
Diagnosing the LLM inference bottleneck (roofline / profiling)Justifying weight and KV cache quantization for faster generationDesigning serving with batching that amortizes weight readsChoosing hardware by memory bandwidth (HBM), not FLOPs aloneMotivating GQA/MQA and speculative decoding

How it works

Compute-kernel performance is described by the roofline model: an operation is compute-bound when its arithmetic intensity (FLOPs/byte) exceeds the hardware's ratio of peak compute to memory bandwidth, and memory-bound otherwise. In the LLM decode phase, at small batch size, each byte of weights read is paired with very few operations (matrix-vector multiply), so intensity is low and memory reads dominate. To generate one token, all model parameters plus the entire KV cache must be read from HBM; the theoretical speed ceiling is (memory bandwidth) / (bytes read per token). Hence techniques that reduce bytes read per token — weight and KV cache quantization, GQA/MQA, batching (amortizing weight reads across requests), and speculative decoding (multiple tokens per pass) — directly raise generation throughput.

Problem solved

The natural assumption that LLM inference is GPU-compute bound leads to wrong optimizations. The memory-bound decoding concept explains why, when generating one token at a time, the GPU is underutilized: matrix-vector operations have low arithmetic intensity (few FLOPs per byte read from memory), so time is dominated by moving weights and the KV cache from HBM. Understanding this steers optimization toward reducing memory traffic (quantization, batching, GQA/MQA, speculative decoding) rather than adding FLOPs.

Key mechanisms

The roofline model and arithmetic intensity (FLOPs/byte)
Reading all weights and the KV cache per token
Amortizing weight reads via batching
Reducing bytes per token: quantization, GQA/MQA
Speculative decoding (multiple tokens per pass)

Strengths & limitations

Strengths
✓Correctly identifies the real LLM inference bottleneck
✓Justifies weight and KV cache quantization as the path to speedup
✓Points to batching as a way to raise arithmetic intensity
✓Guides hardware choice by memory bandwidth (HBM)
✓Motivates GQA/MQA and speculative decoding
Limitations
✗Applies mainly at small batch; at large batch operations may become compute-bound
✗The roofline model is a simplification (ignores cache hierarchy and overlap)
✗It does not remove the limit, only explains and directs around it
✗Reducing bytes per token can cost accuracy
✗For very long contexts, dominance shifts to KV cache traffic

Components

Arithmetic intensityDeterminant of whether an operation is memory- or compute-bound

The ratio of floating-point operations to bytes read from memory; low in the decode phase at small batch.

Roofline modelClassifying the compute bottleneck

An analytical framework relating attainable performance to arithmetic intensity, peak compute, and hardware memory bandwidth.

Memory traffic per tokenThe direct factor limiting decode speed

The total size of weights and KV cache read from HBM to generate one token, setting the theoretical speed limit.

Implementation

Implementation pitfalls
Optimizing FLOPs instead of memory trafficHigh

Focusing on reducing floating-point operations does not speed up decode, because the bottleneck is memory bandwidth.

Fix:Reduce bytes read per token: quantize weights and KV cache, use GQA/MQA, batch requests, and consider speculative decoding.
Ignoring KV cache traffic at long contextMedium

For long contexts, per-token KV cache reads match or exceed weight reads, further loading memory.

Fix:Quantize the KV cache, use GQA/MQA, and trim unnecessary context length.

Evolution

Original paper · 2009 · Communications of the ACM · Samuel Williams
Roofline: An Insightful Visual Performance Model for Multicore Architectures
Samuel Williams, Andrew Waterman, David Patterson
2009
The roofline model formalizes the memory bound
Inflection point

Roofline provides a framework to classify kernels as compute- or memory-bound by arithmetic intensity.

2022
Recognizing LLM decode as memory-bound

LLM serving performance analyses (e.g. works on efficient transformer inference) identify memory-bound decoding as the main limit and the motivation for batching and quantization.

Hyperparameters (configurable axes)

Batch sizeCritical

The main lever for arithmetic intensity; a larger batch amortizes weight reads and eases the memory limit.

1Most strongly memory-bound.
64+Shift toward compute-bound.
Bytes per parameterHigh

Weight precision (FP16/INT8/INT4) setting the bytes read per token.

2 (FP16)Baseline.
~0.5 (INT4)~4x less memory traffic.

Computational complexity

Computational characteristics
→Low arithmetic intensity at small batch
→Speed ~ memory bandwidth / bytes per token
→Memory traffic = model weights + KV cache per token
→Batching shifts operations toward compute-bound
→Dependence on HBM bandwidth, not FLOPs

Time complexity: t_token >= (bajty wag + bajty KV cache) / przepustowość pamięci. Space complexity: Ruch pamieci na token = O(rozmiar wag + rozmiar KV cache).

Benchmark notes

Modern GPUs have a FLOPs-to-memory-bandwidth ratio on the order of hundreds of operations per byte, whereas decode at batch 1 reaches an intensity near ~1-2 operations per byte - hence deep compute underutilization. Moving from FP16 to INT4 weights (~4x fewer bytes per token) yields a roughly proportional several-fold generation speedup, confirming memory-traffic dominance.

Compute bottleneck

Memory bandwidth vs arithmetic intensity

The phenomenon is by definition a memory bottleneck: low arithmetic intensity makes speed depend on memory bandwidth, not FLOPs.

Depends on
Bajty czytane na tokenPrzepustowość pamięci sprzętu

Execution paradigm

Primary mode
Dense

Concerns the dense decode pass in which all weights are read per token.

Activation pattern
All paths active
Routing mechanism

This is a performance property of dense autoregressive generation, not a routing mechanism.

Parallelism

Parallelism level
Partially parallel

Batching many requests amortizes weight reads and eases the memory limit; within a single sequence generation stays sequential.

Scope
InferenceAcross devices

Hardware requirements

Primary

Generation speed depends on GPU memory bandwidth (HBM2e/HBM3); cards with higher bandwidth generate tokens faster regardless of spare FLOPs.

Limited

On CPUs, low RAM bandwidth makes memory-bound decoding even more severe, sharply limiting tokens per second.