Robots Atlas>ROBOTS ATLAS
Inference

LLM Inference

2017ActiveUpdated: 14 August 2026Published
Key innovation
Frames LLM text generation as a two-phase process (prefill + autoregressive decoding) with a KV cache, where the decode phase is memory-bandwidth bound — which defines how latency, throughput and cost are optimized.
Category
Inference
Abstraction level
System
Operation level
InferenceServing
Use cases
Serving chatbots and AI assistantsGenerating text and code in productsBatch processing of promptsEdge/on-device inferenceOptimizing the cost and throughput of LLM deploymentsInference of reasoning models with long 'thinking'

How it works

The prompt is tokenized and processed in the prefill phase: the model computes representations for all tokens at once and stores key-value pairs for the attention layers (the KV cache). Then, in the decode phase, the model uses the KV cache to generate subsequent tokens one at a time; each new token appends its entries to the KV cache. The inference server batches requests, manages KV memory (e.g. PagedAttention), applies sampling parameters (temperature, top-p) and may use acceleration techniques (speculative decoding, quantization, FlashAttention). Generation ends when a stop token or length limit is reached.

Problem solved

Running an LLM to generate text is costly and slow because autoregressive decoding is sequential and memory-bandwidth bound. LLM Inference as a field addresses how to generate tokens fast, cheaply and at high throughput without changing the model weights.

Key mechanisms

Two phases: prefill (parallel) and decode (autoregressive)
The key-value cache (KV cache)
Memory-bandwidth bound decode phase
Continuous batching of requests
Acceleration techniques: PagedAttention, FlashAttention, GQA, speculative decoding, quantization
Sampling parameters (temperature, top-p, top-k)

Strengths & limitations

Strengths
✓Enables practical, real-time use of LLMs
✓A rich toolbox for optimizing latency and throughput
✓Scaling via batching and serving parallelism
✓Flexible quality/cost trade-offs (quantization, reasoning effort)
Limitations
✗The decode phase is sequential and memory-bound — hard to speed up
✗The KV cache grows with context length, consuming GPU memory
✗A latency (small batch) vs throughput (large batch) trade-off
✗High operating cost at scale
✗Long context and reasoning models further increase cost/latency

Components

Prefill phaseContext initialization

Parallel processing of the whole prompt and populating the KV cache; usually compute-bound.

Decode phaseToken generation

Autoregressive token-by-token generation using the KV cache; usually memory-bound.

KV cacheContext memory

A cache of attention key-value pairs, avoiding recomputation for earlier tokens.

Official

Scheduler / batcherServing orchestration

The server layer that batches requests and manages memory (e.g. continuous batching, PagedAttention).

Official

Implementation

Implementation pitfalls
KV cache memory blow-upHigh

Long context and large batches quickly exhaust GPU memory via a growing KV cache.

Fix:PagedAttention, KV quantization, GQA, length limits, offloading.
Poor batch sizing (latency vs throughput)Medium

Large batches raise throughput but hurt single-request latency, and vice versa.

Fix:Continuous batching, prioritization, prefill/decode disaggregation.
Quality loss from aggressive quantizationMedium

Too-low precision (e.g. INT4) can degrade generation quality.

Fix:Calibrated quantization, quality evaluation, mixed precision.

Evolution

2017
Transformer — autoregressive decoding

The architecture underlying token-by-token generation and the KV cache.

2022
FlashAttention — fast, memory-efficient attention

An attention speedup critical for prefill and long context.

2023
PagedAttention / vLLM — continuous batching
Inflection point

Efficient KV-cache management and continuous batching drastically increased serving throughput.

2023
Speculative decoding in practice

A draft model speeds up generation without quality loss.

2024
Prefill/decode disaggregation and serving at scale

Separate resources for prefill and decode phases optimize cost and latency in large deployments.

Hyperparameters (configurable axes)

Batching strategyHigh

How requests are batched; a latency vs throughput trade-off.

continuous batchingDynamically adding requests (e.g. vLLM).
static batchSimpler, worse GPU utilization.
Precision / quantizationHigh

The numeric format of weights/activations affecting speed, memory and quality.

FP16/BF16Default precision.
INT8/FP8/INT4Quantization: less memory, faster, at some quality cost.
Sampling parametersMedium

Temperature, top-p, top-k controlling generation randomness.

temperature, top_p, top_kBalancing determinism vs creativity.

Computational complexity

Time complexity: O(n·d² + n²·d). Space complexity: O(P + L·n·d).

Execution paradigm

Primary mode
Dense

For dense models all weights are active per token; MoE models activate a subset of experts (mixture).

Activation pattern
All paths active

Parallelism

Parallelism level
Partially parallel

Prefill is highly parallel across tokens; decode is sequential across tokens but parallel across requests (batching) and devices (tensor/pipeline parallelism).

Scope
InferenceAcross tokensAcross devices

Hardware requirements

Primary

LLM inference is heavy matrix multiplication; memory bandwidth (HBM) is also critical.

Good fit

Models are also served on TPUs and other accelerators.

Limited

Possible (e.g. llama.cpp) for small models, but slower.