Robots Atlas>ROBOTS ATLAS
Inference

Cross-Model KV Cache Transfer

2026ResearchPublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
Lets a target model reuse the key-value (KV) cache computed by another model from the same family via a learned per-head linear mapping, eliminating costly prefill recomputation.
Category
Inference
Abstraction level
Pattern
Operation level
InferenceServing
Use cases
LLM serving with dynamic model switching (cascading / routing)Reducing time-to-first-token (TTFT) for long contextsMulti-model inference systems sharing contextSame-family model cascades of different sizes

How it works

1. The source model processes the input context and produces a KV cache (key-value states for every layer and attention head). 2. Offline, a closed-form linear mapping (ridge regression) is calibrated per head to transform the source model's KV states into the target model's KV space. 3. Source layers with the highest predictive power for the target layers are selected. 4. Positional encodings are removed so the mapped states are position-independent and reusable. 5. At inference time, instead of running a full prefill, the target model receives the mapped KV states and proceeds directly to decoding. For challenging model pairs, nonlinear variants can improve accuracy.

Problem solved

In production deployments that dynamically switch between differently sized models within one family (e.g. to balance cost and quality), each switch requires the new model to reprocess the entire context (prefill) from scratch. This increases latency (time-to-first-token) and GPU usage. Cross-model KV cache transfer removes this redundancy by porting the already-computed context.

Components

Per-head ridge regression mapperCore of the transfer — converting KV states between models

A closed-form linear mapping (ridge regression) learned separately for each attention head, transforming the source model's KV states into the target model's KV space.

Source layer selectionChoosing the most informative source layers

A mechanism for selecting the source-model layers with the highest predictive power for the target-model layers.

Positional encoding removalEnsuring position independence

Removing positional encoding from the KV states so the mapped states are position-independent and reusable.

Calibration datasetFitting the mapping before deployment

A small dataset used offline to fit the coefficients of the linear mapping.

Implementation

Implementation pitfalls
Limited to one model familyHigh

The transfer works for models sharing architecture and tokenizer within a family; it does not port freely between unrelated models.

Fix:Use only within a verified model family and calibrate the mapper for the specific source–target pair.
Accuracy loss on challenging pairsMedium

For some model pairs accuracy drops below the full 73–98% range; the linear mapping may be insufficient.

Fix:Apply nonlinear mapping variants for challenging model pairs.
Handling positional encodingMedium

To make mapped states reusable, positional encoding must be removed, which requires compatibility with the model's positional scheme.

Fix:Strip positional encoding before mapping and reapply the target model's positional scheme.

Evolution

Original paper · 2026 · arXiv preprint · Taekyung Heo
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, Mohammad Mahdi Kamani, Ritika Borkar, Makesh Tarun Chandran, Pantea Zardoshti, Bita Darvish Rouhani
2025
Cross-model prefix KV cache reuse between base and adapted models (Activated LoRA)

First LLM serving engine supporting cross-model prefix KV cache reuse between a base model and its adapters.

2026
Closed-form linear mapping for cross-model KV transfer within a family
Inflection point

Discovery of linear structure in cross-model KV pairs and introduction of a closed-form per-head mapping that replaces full prefill.

Parallelism

Parallelism level
Fully parallel

The linear mapping is applied independently per attention head and layer, so the transfer is fully parallelizable — unlike sequential prefill.

Scope
Inference

Hardware requirements

Primary

The technique targets GPU-based LLM serving; the mapping itself is a lightweight linear operation (matrix multiply) well suited to tensor cores.