Decode
How it works
In the decode phase the model runs one forward pass on a single token (the most recently generated one). In the attention layers the new token computes its query vector (Q) and attends over the keys and values of all previous tokens fetched from the KV cache; its own K, V are appended to the cache. The remaining layers (MLP, projections) process a single vector. The model produces the next-token distribution, samples it, and repeats until an end token or length limit. Because operations concern a single token (small matrix-vector multiplications), the compute units are underutilized and step time is dominated by reading weights and the KV cache from memory (memory-bound). Batching many requests and larger KV caches increase arithmetic intensity and throughput.
Problem solved
After the context is built in prefill, the model must generate subsequent tokens sequentially because each token depends on the previous ones. The decode phase does this efficiently: instead of recomputing attention over the whole sequence, it uses the stored KV cache and processes only one new token per step. Because each token requires reading all model weights and the entire KV cache from memory, for a single request the phase is memory-bandwidth bound rather than compute bound.
Key mechanisms
Strengths & limitations
Components
A forward pass on a single token producing the next-token distribution; matrix-vector operations with low arithmetic intensity.
The new token attends over all stored keys and values; its own K, V are appended to the cache each step.
Selecting the next token from the distribution (greedy, top-k, top-p, temperature) and feeding it as input to the next step.
Implementation
With one or a few requests, matrix-vector operations barely load the compute units, wasting the GPU's FLOPs potential.
The KV cache grows linearly with sequence length and request count, becoming a memory bottleneck for long contexts.
Evolution
The spread of generative transformers exposes the memory-bound nature of decode and the role of the KV cache.
Continuous batching (Orca, vLLM) and speculative decoding raise throughput and cut decode-phase latency.
Hyperparameters (configurable axes)
Number of concurrent requests; larger amortizes weight reads and raises throughput.
Token-selection strategy and parameters: temperature, top-k, top-p.
Computational complexity
Time complexity: O(N d) na token (uwaga po KV cache) + O(d^2) projekcje. Space complexity: O((N + t) x L x d) na rosnacy KV cache.
For a single request the theoretical decode-speed ceiling is (memory bandwidth) / (weight size + KV cache per token); e.g. a 13B model in FP16 (~26 GB) on a ~1 TB/s card yields on the order of tens of tokens/s. Continuous batching can multiply the system's total throughput by amortizing weight reads across many concurrent requests.
Compute bottleneck
Each decode step must read all model weights and the entire KV cache from memory while doing few operations, so memory bandwidth sets the speed.
Execution paradigm
All layers and compute paths are active for each generated token.
Decode runs a dense forward pass on a single token with no conditional routing (unless the model is MoE).
Parallelism
Tokens are generated sequentially (each depends on the previous); parallelism comes across requests via batching, not within a single sequence.
Hardware requirements
GPUs with high memory bandwidth (HBM) are preferred because decode is memory-bound; memory bandwidth, not FLOPs, sets generation speed.
Decode runs on CPU (e.g. llama.cpp), but low RAM bandwidth limits tokens per second.