Robots Atlas>ROBOTS ATLAS
Inference

Inference Latency

2023ActiveUpdated: 14 August 2026Published
Key innovation
Decomposes a model's response delay into measurable components (TTFT, TPOT/ITL, end-to-end latency), enabling deliberate optimization of user experience separately from throughput.
Category
Inference
Abstraction level
Primitive
Operation level
InferenceServing
Use cases
Defining SLAs and responsiveness targets for LLM productsComparing models and providers (latency benchmarks)Choosing serving optimization techniquesDesigning interactive and voice assistantsBalancing the latency vs throughput vs cost trade-off

How it works

Latency is measured along a single request's timeline: from sending the prompt to the first token (TTFT), then the gaps between subsequent tokens (ITL) and the moment of the last token (end-to-end latency). TTFT depends mainly on the prefill phase (prompt length, compute), while TPOT depends on the decode phase (memory-bandwidth bound: reading weights and the KV cache at each step). Optimization means shrinking these components: a smaller/quantized model, GQA and a smaller KV cache, FlashAttention, speculative decoding, better hardware (HBM), prefill/decode disaggregation, and choosing a batch size that meets the latency target.

Problem solved

Slow responses hurt user experience and limit interactive and agentic use. Inference Latency as a metric makes it possible to precisely measure and optimize responsiveness (TTFT, TPOT) independently of throughput.

Key mechanisms

TTFT (Time To First Token)
TPOT / ITL (Time Per Output Token / Inter-Token Latency)
End-to-end latency โ‰ˆ TTFT + number_of_tokens ร— TPOT
Dependence on memory bandwidth (the decode phase)
A trade-off with throughput (batch size)
The impact of context length, model size and hardware

Strengths & limitations

Strengths
โœ“Directly measures user-perceived responsiveness
โœ“Decomposition into components (TTFT/TPOT) shows what to optimize
โœ“Enables SLAs and cross-system comparisons
โœ“Complements throughput and cost for a full performance picture
Limitations
โœ—Often in conflict with throughput and cost (a trade-off)
โœ—Depends on measurement context (prompt/response length, load)
โœ—Variability (tail, p99) is harder to handle than the mean
โœ—Grows with long context and reasoning models
โœ—Hard to maintain under high, variable traffic

Components

TTFT (Time To First Token)Startup latency

Time from sending the prompt to the first generated token; dominated by the prefill phase.

TPOT / ITLStreaming latency

Average time to generate each subsequent token during decoding.

End-to-end latencyTotal latency

Total time to complete the response (โ‰ˆ TTFT + number_of_tokens ร— TPOT).

Implementation

Implementation pitfalls
Reporting only the mean, not the tail (p99)Medium

Mean latency hides bad experiences; the tail (p95/p99) is what matters.

Fix:Measure percentiles (p50/p95/p99) under realistic load.
Confusing latency with throughputMedium

Optimizing throughput (large batch) can worsen user latency.

Fix:Separate targets for latency and throughput; continuous batching, priorities.
Measuring outside realistic conditionsLow

Measurements without variable traffic, context length and network don't reflect production.

Fix:Benchmark on representative workloads and lengths.

Evolution

2022
Latency as a key LLM serving metric

With production LLM use, latency becomes a first-class metric alongside throughput.

2023
Standardization of TTFT and TPOT/ITL
Inflection point

TTFT and time-per-token become standard metrics for LLM comparisons and SLAs.

2025
Latency of reasoning models

Long 'thinking' chains materially increase end-to-end latency, forcing new trade-offs.

Hyperparameters (configurable axes)

Batch sizeHigh

A larger batch raises throughput but hurts single-request latency.

smallLow latency, lower throughput.
largeHigh throughput, higher latency.
Model size / precisionHigh

A smaller or quantized model reduces latency at some quality cost.

quantized (INT8/FP8)Faster decoding, less memory.
Context lengthMedium

A longer prompt lengthens prefill (TTFT), and a growing KV cache raises TPOT.

short vs longLong context = higher latency.

Computational complexity

Time complexity: L โ‰ˆ TTFT + nยทTPOT.

Hardware requirements

Primary

Decode latency is memory-bound โ€” high GPU memory bandwidth (HBM) is key.

Good fit

High-memory-bandwidth accelerators also lower TPOT.