Robots Atlas>ROBOTS ATLAS
Inference

Prefill

2020ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Isolating the first phase of LLM inference, in which the entire prompt is processed in parallel in a single forward pass, filling the key-value cache (KV cache) and producing the first token — a compute-bound phase, the opposite of the decode phase.
Category
Inference
Abstraction level
Pattern
Operation level
InferenceServing
Use cases
Processing the prompt and system context before generationMeasuring and optimizing time to first token (TTFT)Chunked prefill to interleave long prompts with the decode phasePrefix caching (sharing KV cache of common prefixes)Splitting prefill and decode onto separate nodes (disaggregated serving)

How it works

In the prefill phase the model runs one (or a few, with chunked prefill) forward pass over the entire prompt sequence of length N. Because all input tokens are known in advance, the attention and MLP computations are done as large matrix multiplications with high unit utilization (compute-bound). During this pass, for each layer the keys and values (K, V) of all N tokens are written to the KV cache. At the end the model produces a probability distribution over the next token and samples the first generated token. Control then moves to the decode phase, which generates subsequent tokens one at a time using the filled KV cache. The compute cost of prefill grows with prompt length (attention ~O(N^2)).

Problem solved

Autoregressive generation requires the model to process the entire prompt and build a context representation before producing the first token. The prefill phase does this efficiently by processing all prompt tokens at once (in parallel) rather than one by one, and by writing the attention keys and values into the KV cache so the decode phase need not recompute them. Prefill determines the time to first token and heavily loads the compute units.

Key mechanisms

Parallel forward pass over the entire prompt sequence
Matrix-matrix multiplications with high arithmetic intensity
Populating the KV cache with keys and values of all tokens
First-token generation and TTFT determination
Chunked prefill interleaving fragments with decode steps

Strengths & limitations

Strengths
✓Parallel processing of the whole prompt in a single pass
✓High compute-unit utilization (compute-bound)
✓Builds the KV cache, eliminating context recomputation in decode
✓Amenable to prefix caching of common prefixes (system prompt)
✓Can be separated from decode (disaggregated serving)
Limitations
✗Cost grows quadratically with prompt length (attention O(N^2))
✗A long prefill can block other requests' generation in the batch
✗Determines TTFT, which can be high for long prompts
✗Requires large KV cache memory at long context
✗Without prefix caching, repeated prefixes are recomputed from scratch

Components

Parallel prompt forward passBuilding the context representation

Processing all N prompt tokens simultaneously in a single pass, realized as large matrix multiplications.

KV cache populationPreparing the cache for the decode phase

Writing the attention keys and values of all prompt tokens for every layer, to avoid recomputing them in the decode phase.

First token generationTransition to the autoregressive phase

Computing the next-token distribution at the last position and sampling the first output, which determines TTFT.

Implementation

Implementation pitfalls
Long prompt blocking other requests' generationMedium

A single long prefill can occupy the GPU for a long time and increase token latencies of other requests in the batch.

Fix:Use chunked prefill to interleave prefill chunks with decode steps and keep generation smooth.
Missed prefix cachingLow

Repeated prefixes (e.g. a shared system prompt) recomputed from scratch waste prefill compute.

Fix:Enable prefix caching / automatic prefix caching to share the KV cache of common prefixes.

Evolution

Original paper · 2023 · SOSP 2023 · Woosuk Kwon
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Ion Stoica
2020
Prefill vs decode phases distinguished in autoregressive serving
Inflection point

As generative transformers spread, inference splits visibly into compute-bound prefill and memory-bound decode.

2023
Chunked prefill and disaggregated serving

Work on efficient serving (e.g. vLLM, Sarathi, DistServe) introduces chunked prefill and prefill/decode disaggregation for better GPU utilization.

Hyperparameters (configurable axes)

Chunk size (chunked prefill)High

Number of prompt tokens processed per chunk when interleaving with decode.

512A typical chunk size.
Prefix cachingMedium

Sharing the KV cache of common prefixes across requests.

enabledReduces prefill time for repeated prefixes.

Computational complexity

Computational characteristics
→Compute-bound phase (high FLOPs utilization)
→Attention cost ~O(N^2) with respect to prompt length
→High arithmetic intensity (matrix-matrix multiplications)
→Writes the full prompt KV cache to memory
→Determines TTFT (time to first token)

Time complexity: O(N^2 d) uwaga + O(N d^2) projekcje. Space complexity: O(N x L x d) na KV cache promptu.

Benchmark notes

TTFT dominated by prefill grows with prompt length; for long contexts the O(N^2) attention term becomes noticeable. Prefix caching of a shared system prompt can cut prefill time to nearly zero for repeated prefixes, while chunked prefill smooths other requests' token latencies by interleaving prefill fragments with decode steps.

Compute bottleneck

Matrix multiplication (compute-bound)

Prefill is compute-bound: the large matrix-matrix multiplications of attention and MLP saturate Tensor Core units.

Depends on
Długość promptu NPrzepustowość obliczeniowa GPU

Execution paradigm

Primary mode
Dense

All tokens and compute paths are active simultaneously.

Activation pattern
All paths active
Routing mechanism

Prefill runs a dense forward pass with no conditional routing (unless the model is MoE).

Parallelism

Parallelism level
Fully parallel

All prompt tokens are processed simultaneously, making prefill highly parallel and compute-bound.

Scope
InferenceAcross tokens

Hardware requirements

Primary

The large matrix multiplications of prefill maximally utilize Tensor Cores; this is the phase where the GPU reaches high FLOPs utilization.

Good fit

The parallel, compute-bound nature of prefill maps well to the MXU units in TPUs.