Robots Atlas>ROBOTS ATLAS
Architecture

QK-Norm

2020ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
Normalizes the query (Q) and key (K) vectors before the attention dot product and replaces the fixed 1/√d scaling with a learnable parameter, so attention logits do not diverge and the softmax avoids saturation.
Category
Architecture
Abstraction level
Building block
Operation level
Architecture blockLayerTraining
Use cases
Stabilizing training of large transformersLow-resource machine translationScaling Vision TransformersLarge language models (LLMs)Preventing attention softmax saturation

How it works

1) After the linear projections to Q and K, each query and key vector is normalized along the head dimension (head_dim). The original uses L2 normalization; variants use LayerNorm (ViT-22B) or RMSNorm (Qwen3). 2) The normalized Q and K are multiplied (dot product). After L2 normalization the dot product is bounded to [-1, 1]. 3) Instead of dividing by √d, the result is scaled by a learnable parameter (temperature g) that controls the sharpness of the softmax. 4) The rest of attention (softmax, multiply by V) is unchanged.

Problem solved

Without normalization, attention logits (Q·K dot products) can grow to very large magnitudes as models scale, causing softmax saturation: attention weights collapse toward near one-hot, near-zero-entropy distributions, gradients vanish, and training of large models diverges. QK-Norm bounds the logit range and stabilizes training.

Components

Q/K normalizationBounding the attention-logit range

A step that normalizes query and key vectors along the head dimension before the dot product.

INQ and K vectors after the linear projection.
OUTNormalized Q and K.
L2 normalizationOriginal QKNorm (Henry et al. 2020): L2 normalization along the head dimension.
QK LayerNormViT-22B variant: LayerNorm on Q and K.
QK RMSNormQwen3 variant: RMSNorm on Q and K.

Official

Learnable scaling parameter (temperature g)Controls softmax sharpness

A learnable parameter replacing the fixed 1/√d scaling; controls the sharpness of the softmax over the normalized Q·K product.

Implementation

Implementation pitfalls
Omitting the learnable scaleHigh

After L2 normalization the Q·K product is bounded to [-1,1]; without the learnable parameter g the logits are too small and the softmax too flat.

Fix:Always add a learnable scaling parameter (temperature) instead of 1/√d.
Wrong axis or normalization typeMedium

Normalization must be per-head along head_dim; the choice of variant (L2 vs LayerNorm vs RMSNorm) affects stability.

Fix:Normalize along head_dim and pick the variant consistent with the reference architecture.
Ordering relative to RoPEMedium

QK-Norm and rotary position embeddings (RoPE) operate on the same Q/K; the order of application affects the result.

Fix:Fix and consistently apply the order (normalization vs RoPE) matching the reference implementation.

Evolution

Original paper · 2020 · Findings of EMNLP 2020 · Alex Henry
Query-Key Normalization for Transformers
Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, Yuxuan Chen
2020
QKNorm introduced
Inflection point

Henry et al. propose QKNorm: L2 normalization of Q and K along the head dimension plus scaling by a learnable parameter instead of 1/√d; average BLEU gain of 0.928 on 5 low-resource pairs.

2023
Adoption in ViT-22B (QK LayerNorm)
Inflection point

Google Research applies LayerNorm to Q and K in ViT-22B to prevent attention-logit divergence and near one-hot, zero-entropy attention distributions in 8B+ models.

2025
QK-Norm (RMSNorm) in Qwen3

Qwen3 removes the QKV bias from Qwen2 and introduces QK-Norm (RMSNorm) into the attention mechanism to ensure stable training.

Hyperparameters (configurable axes)

Normalization typeHigh

L2 (original), LayerNorm (ViT-22B), or RMSNorm (Qwen3).

L2
LayerNorm
RMSNorm
Learnable scale / temperatureCritical

Parameter g scaling the Q·K product instead of 1/√d.

Normalization axisMedium

Per-head normalization along head_dim.

head_dim

Computational complexity

Time complexity: O(n·d) dodatkowo.

Execution paradigm

Primary mode
Dense

A modification of dense attention; all heads and paths remain active.

Activation pattern
All paths active

Parallelism

Parallelism level
Fully parallel

Normalization is independent per token and per head, so it is fully parallel.

Scope
TrainingInferenceAcross tokens

Hardware requirements

Good fit

A pointwise normalization runs on any accelerator without special requirements.

Primary

Transformers with QK-Norm are trained mainly on GPUs with Tensor Cores.