Robots Atlas>ROBOTS ATLAS
Inference

Decode

2020ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Isolating the autoregressive phase of LLM inference, in which tokens are generated one at a time โ€” each step reads the entire KV cache and all model weights to produce a single token, making the phase memory-bandwidth bound (memory-bound), the opposite of compute-bound prefill.
Category
Inference
Abstraction level
Pattern
Operation level
InferenceServing
Use cases
Autoregressive text generation token by tokenMeasuring inter-token latency (ITL) and throughputContinuous batching of many requests for better GPU utilizationSpeculative decoding to accelerate the decode phasePlacing decode on separate nodes (disaggregated serving)

How it works

In the decode phase the model runs one forward pass on a single token (the most recently generated one). In the attention layers the new token computes its query vector (Q) and attends over the keys and values of all previous tokens fetched from the KV cache; its own K, V are appended to the cache. The remaining layers (MLP, projections) process a single vector. The model produces the next-token distribution, samples it, and repeats until an end token or length limit. Because operations concern a single token (small matrix-vector multiplications), the compute units are underutilized and step time is dominated by reading weights and the KV cache from memory (memory-bound). Batching many requests and larger KV caches increase arithmetic intensity and throughput.

Problem solved

After the context is built in prefill, the model must generate subsequent tokens sequentially because each token depends on the previous ones. The decode phase does this efficiently: instead of recomputing attention over the whole sequence, it uses the stored KV cache and processes only one new token per step. Because each token requires reading all model weights and the entire KV cache from memory, for a single request the phase is memory-bandwidth bound rather than compute bound.

Key mechanisms

Autoregressive step on a single token (matrix-vector)
Attention over the KV cache and appending new K, V
Token sampling (greedy, top-k, top-p, temperature)
Continuous batching of many requests
Speculative decoding (multiple tokens per verification pass)

Strengths & limitations

Strengths
โœ“Uses the KV cache, avoiding context recomputation
โœ“Token-by-token generation with full sampling control
โœ“Amortizes weight reads via continuous batching
โœ“Amenable to speculative-decoding acceleration
โœ“Readily optimized by weight and KV cache quantization
Limitations
โœ—Memory-bound phase: speed limited by memory bandwidth
โœ—GPU underutilization at small batch (matrix-vector)
โœ—Inherently sequential (each token depends on the previous)
โœ—KV cache grows linearly with length and request count
โœ—Long context increases per-token memory traffic

Components

Autoregressive step (single token)Producing one next token

A forward pass on a single token producing the next-token distribution; matrix-vector operations with low arithmetic intensity.

KV cache read and updateReusing context without recomputation

The new token attends over all stored keys and values; its own K, V are appended to the cache each step.

Token samplingControlling generation and randomness

Selecting the next token from the distribution (greedy, top-k, top-p, temperature) and feeding it as input to the next step.

Implementation

Implementation pitfalls
GPU underutilization at small batchHigh

With one or a few requests, matrix-vector operations barely load the compute units, wasting the GPU's FLOPs potential.

Fix:Use continuous batching to combine many requests and raise arithmetic intensity and throughput.
KV cache growth with long contextMedium

The KV cache grows linearly with sequence length and request count, becoming a memory bottleneck for long contexts.

Fix:Use PagedAttention, KV cache quantization, or grouped attention (GQA/MQA) to limit memory usage.

Evolution

Original paper ยท 2023 ยท SOSP 2023 ยท Woosuk Kwon
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Ion Stoica
2020
Prefill vs decode phases distinguished in autoregressive serving
Inflection point

The spread of generative transformers exposes the memory-bound nature of decode and the role of the KV cache.

2023
Continuous batching and speculative decoding

Continuous batching (Orca, vLLM) and speculative decoding raise throughput and cut decode-phase latency.

Hyperparameters (configurable axes)

Batch sizeHigh

Number of concurrent requests; larger amortizes weight reads and raises throughput.

1Lowest arithmetic intensity, memory-bound.
32-256High-throughput serving (continuous batching).
Sampling parametersMedium

Token-selection strategy and parameters: temperature, top-k, top-p.

temperature=0Deterministic (greedy).
top-p=0.9Nucleus sampling.

Computational complexity

Computational characteristics
โ†’Memory-bound phase (low arithmetic intensity)
โ†’One token per step, matrix-vector operations
โ†’Reads all weights and the KV cache per token
โ†’Determines ITL (inter-token latency) and tokens/s
โ†’Throughput increases with batching

Time complexity: O(N d) na token (uwaga po KV cache) + O(d^2) projekcje. Space complexity: O((N + t) x L x d) na rosnacy KV cache.

Benchmark notes

For a single request the theoretical decode-speed ceiling is (memory bandwidth) / (weight size + KV cache per token); e.g. a 13B model in FP16 (~26 GB) on a ~1 TB/s card yields on the order of tens of tokens/s. Continuous batching can multiply the system's total throughput by amortizing weight reads across many concurrent requests.

Compute bottleneck

Memory bandwidth (memory-bound)

Each decode step must read all model weights and the entire KV cache from memory while doing few operations, so memory bandwidth sets the speed.

Depends on
Rozmiar wag i KV cacheRozmiar batcha

Execution paradigm

Primary mode
Dense

All layers and compute paths are active for each generated token.

Activation pattern
All paths active
Routing mechanism

Decode runs a dense forward pass on a single token with no conditional routing (unless the model is MoE).

Parallelism

Parallelism level
Sequential

Tokens are generated sequentially (each depends on the previous); parallelism comes across requests via batching, not within a single sequence.

Scope
InferenceAcross devices

Hardware requirements

Primary

GPUs with high memory bandwidth (HBM) are preferred because decode is memory-bound; memory bandwidth, not FLOPs, sets generation speed.

Possible

Decode runs on CPU (e.g. llama.cpp), but low RAM bandwidth limits tokens per second.