Robots Atlas>ROBOTS ATLAS
Inference

Re-prefill

ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
It is not an optimization — it is the baseline, „naive" operation of recomputing the entire KV cache from scratch, the reference point that many caching and KV-transfer techniques aim to eliminate.
Category
Inference
Abstraction level
Building block
Operation level
InferenceServing
Use cases
Rebuilding the cache after eviction from GPU memoryFallback after a context edit in an agentic loopReprocessing context after a prompt prefix changeA comparison baseline for KV caching and transfer techniques

How it works

1. The model takes the entire (or entirely changed) input context. 2. It runs a full forward pass through all layers, recomputing keys (K) and values (V) for every token and attention head. 3. The freshly computed states repopulate the KV cache. 4. The model emits the first token and moves to the decoding phase using the reconstructed cache. Because this is a full prefill, attention cost scales quadratically with context length, and the whole pass is repeated whenever the prior cache was invalidated.

Problem solved

When the KV cache becomes stale or unavailable (memory eviction, context edit, prefix change, model switch), decoding cannot continue from an incorrect or missing cache. Re-prefill guarantees correctness by reconstructing the KV states from scratch — at the cost of time and compute.

Implementation

Implementation pitfalls
Full recomputation cost on every invalidationHigh

Every context edit or cache eviction forces a full re-prefill, significantly raising time-to-first-token (TTFT).

Fix:Use prefix caching, KV cache reuse/transfer, and stabilize prompt prefixes to avoid invalidation.
Quadratic scaling with context lengthHigh

Attention cost in prefill grows as O(n²), so re-prefilling a long context is especially expensive.

Fix:Apply chunked prefill and context compression/reduction techniques.

Evolution

2026
Cross-model KV cache transfer as an alternative to re-prefill

It was shown that re-prefill after a model switch can be skipped by transferring the KV cache between same-family models — 2.7–25× faster than re-prefill.

Computational complexity

Time complexity: O(n² · d). Space complexity: O(n · d · L).

Execution paradigm

Primary mode
Dense

Prefill is dense compute: all context tokens pass through all layers in a single pass.

Activation pattern
All paths active

Parallelism

Parallelism level
Fully parallel

Unlike sequential decoding, prefill (and thus re-prefill) processes all context tokens in parallel in a single forward pass.

Scope
InferenceAcross tokens

Hardware requirements

Primary

Re-prefill is compute-bound and dominated by matrix multiplications (GEMM), which map ideally onto GPU tensor cores.