Robots Atlas>ROBOTS ATLAS
Inference

TTFT

ActivePublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
Isolating first-token latency as a separate service-level objective, decoupled from throughput and per-token generation time.
Category
Inference
Abstraction level
Primitive
Operation level
InferenceServing
Use cases
Benchmarking and profiling LLM inference serversDefining SLOs/SLAs for LLM servicesOptimising UX of streaming apps and chatbotsAutoscaling and inference capacity planningComparing hardware and inference engines (vLLM, TensorRT-LLM)Latency-sensitive voice and agentic applications

How it works

TTFT is measured as the elapsed time between sending a request and receiving the first non-empty token (this requires streaming mode). The total time comprises: (1) queueing and admission of the request into a batch, (2) the prefill phase — processing all prompt tokens in parallel across all model layers while building the KV cache, (3) computing the logits at the last position and sampling the first token, and (4) detokenisation and network transfer. The prefill phase is compute-bound (a matrix-matrix operation that saturates the GPU), so TTFT grows with prompt length and model size. Common techniques to reduce TTFT include chunked-prefill, prefix/KV cache reuse, prefill-decode disaggregation, tensor parallelism and more powerful compute hardware.

Problem solved

Throughput (tokens/s) and total latency alone do not capture how quickly a user sees the start of a response. TTFT isolates the delay before generation begins — the factor that, in streaming and conversational interfaces, determines perceived responsiveness — and lets teams optimise and set SLOs for the prefill phase independently of the decode phase.

Components

Request queueingAdds load-dependent overhead included in end-to-end TTFT.

Time a request waits to be admitted into the current batch by the inference server scheduler before computation begins.

Prefill phase (prompt processing)Dominates TTFT and grows with prompt length and model size.

Parallel processing of all prompt tokens across all model layers; the main, compute-bound component of TTFT.

INB prompt sequences of n tokens each.
OUTPopulated KV cache for all layers and logits at the last position.
KV cache populationA side effect of prefill; drives memory usage and enables prefix caching that shortens TTFT.

Storing keys and values (K, V) for all prompt tokens and layers, later reused during the decode phase.

First-token samplingConcludes the TTFT measurement — the moment the first token is emitted to the user.

Decoding the first output token from the last-position logits (greedy, top-k, top-p, temperature).

Implementation

Implementation pitfalls
Measuring TTFT without streamingHigh

Without streaming the server returns the whole response at once, so real time-to-first-token cannot be measured.

Fix:Always measure TTFT in streaming mode and record the timestamp of the first non-empty chunk.
Ambiguous definition: with vs without queueingMedium

Pure TTFT (from prefill start) and end-to-end TTFT (including queueing) yield different numbers, hampering comparisons.

Fix:Clearly document whether TTFT includes queueing and compare metrics measured the same way.
Cold start / warm-upMedium

Weight loading, graph compilation and memory allocation inflate the first TTFT measurement.

Fix:Run warm-up requests and exclude them from statistics.
First empty/role chunk instead of first tokenMedium

Some APIs return an empty or metadata first chunk (e.g. a role), which understates TTFT if counted as a token.

Fix:Count TTFT only from the first chunk that contains token content.
Client-side tokenisation and network overheadLow

Network and detokenisation delays affect client-perceived TTFT even though server-side measurements often omit them.

Fix:Measure TTFT as close to the client as possible and report the measurement point.

Evolution

2023
TTFT standardised as a core LLM serving metric

TTFT, together with TPOT and throughput, becomes a standard set of metrics in LLM inference performance engineering (e.g. Databricks and NVIDIA guides).

2023
Splitwise: prefill/decode phase splitting

Splitting prompt computation and token generation across separate machines, enabling independent management of TTFT and generation throughput.

2024
DistServe: prefill/decode disaggregation for separate TTFT and TPOT SLOs
Inflection point

Disaggregating prefill and decode onto different GPUs allows independent optimisation of TTFT (prefill) and per-token time (decode) instead of trading them off.

2024
Sarathi-Serve: chunked-prefill balancing TTFT and throughput

Splitting prefill into chunks with stall-free scheduling sustains high throughput while controlling the impact of batching on TTFT and latency.

Hyperparameters (configurable axes)

Prompt / input lengthCritical

Number of input tokens; the main driver of TTFT.

Model sizeHigh

Parameter and layer count affects prefill cost per token.

Compute hardware (FLOPS / tensor cores)High

Accelerator compute throughput determines prefill speed.

Batch size / concurrencyHigh

Larger batches and load increase queueing and can raise TTFT.

Prefix / KV cache reuseHigh

Caching a shared prefix shortens prefill and TTFT.

Chunked-prefill chunk sizeMedium

Splitting prefill into chunks balances TTFT against throughput.

Tensor / pipeline parallelismMedium

Sharding the model across GPUs can shorten prefill for large models.

Computational complexity

Time complexity: O(n^2 * d). Space complexity: O(n * d * L).

Compute bottleneck

Prefill phase (prompt processing)

TTFT is dominated by the compute-bound prefill — a dense matrix-matrix operation that saturates the accelerator compute units.

Depends on
Długość promptuRozmiar modeluMoc obliczeniowa (FLOPS / tensor cores)Kolejkowanie i batching

Execution paradigm

Primary mode
Dense

TTFT measures the prefill cost — a dense matrix-matrix operation activating all model parameters over all prompt tokens.

Activation pattern
All paths active

Parallelism

Parallelism level
Fully parallel

Prefill processes all prompt tokens in parallel (matrix-matrix), unlike the sequential decode phase, which TTFT does not include.

Scope
InferenceAcross tokens

Hardware requirements

Primary

Prefill is a compute-bound matrix-matrix operation that benefits strongly from high FLOPS and tensor cores, shortening TTFT.

Good fit

Matrix accelerators (TPUs) handle compute-bound prefill for large prompts well.

Limited

On CPUs, prefill of long prompts is slow, leading to high TTFT.