Robots Atlas>ROBOTS ATLAS
Architecture

Cross-Attention

2017ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
An attention variant where queries (Q) come from one sequence while keys and values (K, V) come from another, letting one stream condition on another (e.g. decoder on encoder, image on text).
Category
Architecture
Abstraction level
Building block
Operation level
Architecture blockLayer
Use cases
Encoder-decoder attention in the Transformer (machine translation, summarization)Text conditioning in Latent Diffusion / Stable DiffusionMultimodal models (text-image fusion, e.g. Flamingo)Perceiver / Perceiver IO (projection onto a latent array)Vision-Language-Action models (fusing visual observation with language)

How it works

1. The "querying" sequence (e.g. the decoder) is linearly projected to a query matrix Q. 2. The "source" sequence (e.g. the encoder output) is projected to key K and value V matrices. 3. Dot products Q·Kᵀ are computed, scaled by 1/√d_k and passed through softmax, producing attention weights that show how strongly each querying position attends to each source position. 4. These weights multiply the values V, yielding context representations. 5. In practice the operation is multi-head: Q, K, V are split into h heads computed in parallel, then concatenated and projected by an output matrix. Key point: Q has a different source than K and V, the sequences may have different lengths, and no causal mask is applied (unlike the decoder's masked self-attention). During autoregressive decoding the encoder's K and V are computed once and cached across all steps.

Problem solved

How can one sequence (or modality) draw on information contained in another, possibly of a different length? Cross-attention solves conditional fusion of two representations — e.g. aligning a translation with its source sentence, an image with a text caption, or a robot action with a visual observation — without compressing the source sequence into a single vector, which was the bottleneck of earlier encoder-decoder architectures.

Components

Query projection (from the querying sequence)Projects the target/decoder sequence into the query space Q.

A linear projection W_Q applied to the querying sequence. This is what distinguishes cross-attention from self-attention — Q comes from a different stream than K and V.

INQuerying sequence (e.g. decoder states).
OUTQueries split into heads.
Key & Value projections (from the source sequence)Project the source/encoder sequence into keys K and values V.

Projections W_K and W_V applied to the source sequence. In autoregressive decoding they are computed once and cached, since the source does not change.

INSource sequence (e.g. encoder output).
OUTKeys and values split into heads.
Scaled dot-product attention + softmaxComputes attention weights between querying and source positions and aggregates values.

softmax(Q·Kᵀ / √d_k)·V. No causal mask; an optional padding mask on the source side.

OUTPer-head context representations.

Official

Multi-head concat + output projectionConcatenates head outputs and projects back to d_model.

Concatenation of the h heads and a linear projection W_O.

OUTOutput of the cross-attention layer.

Implementation

Implementation pitfalls
Swapping the source of Q with K/VHigh

Q must come from the querying sequence (decoder) and K, V from the source sequence (encoder). Swapping the sources breaks the conditioning.

Fix:Explicitly pass separate query and key/value tensors to the attention layer and assert their shapes.
Incorrect maskingMedium

Cross-attention uses no causal mask — only a source-side padding mask is needed. Adding a causal mask wrongly restricts access to the input.

Fix:Apply only a key_padding_mask for the source; do not pass a causal mask.
Not caching encoder K/VMedium

In autoregressive decoding the source K and V do not change; recomputing them at every step wastes compute.

Fix:Compute encoder K and V once and store them in a cross-attention cache for the whole decoding run.

Evolution

Original paper · 2017 · NeurIPS 2017 · Ashish Vaswani
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin
2014
Attention in RNN encoder-decoder (precursor of cross-attention)
Inflection point

Bahdanau et al. introduce attention between the decoder and encoder states in machine translation — the decoder attends over the source sequence.

2017
Transformer formalizes encoder-decoder (cross) attention
Inflection point

Vaswani et al. define the encoder-decoder attention layer with scaled dot-product attention and multiple heads; Q from the decoder, K and V from the encoder.

2021
Perceiver — cross-attention to a latent array

Perceiver uses cross-attention to project very large inputs onto a compact latent array, decoupling cost from input length.

2021
Latent Diffusion — text conditioning via cross-attention
Inflection point

Rombach et al. add cross-attention to the UNet to condition image generation on text embeddings — the basis of Stable Diffusion.

2022
Flamingo — gated cross-attention in a multimodal model

Flamingo injects visual representations into a frozen language model through gated cross-attention layers.

Hyperparameters (configurable axes)

Number of heads (h)High

Number of parallel attention heads.

8Transformer base (Vaswani et al., 2017).
Key/query dimension (d_k)Medium

Per-head dimension used in the 1/√d_k scaling.

64d_model 512 / 8 heads.
Model dimension (d_model)Medium

Dimension of the layer's input and output embeddings.

512Transformer base.

Computational complexity

Time complexity: O(n_q · n_kv · d). Space complexity: O(n_q · n_kv).

Compute bottleneck

Q·Kᵀ matmul and softmax

The main cost is the dot products between queries and keys plus softmax normalization; with long source sequences it grows linearly with n_kv.

Execution paradigm

Primary mode
Dense

Dense attention: every querying position attends to all source positions (except masked padding).

Activation pattern
All paths active
Routing mechanism

Parallelism

Parallelism level
Fully parallel

In training all querying positions are computed in parallel. In autoregressive decoding querying positions are produced sequentially, but source K and V are computed once and cached.

Scope
TrainingAcross tokens

Hardware requirements

Primary

The operation is dominated by dense matrix multiplications (Q·Kᵀ, ·V), ideal for tensor cores.

Good fit

Systolic matrix-multiply units handle batched attention products well.

Possible

Runs on CPU and other accelerators, but with lower throughput for long source sequences.