Robots Atlas>ROBOTS ATLAS
Architecture

RMSNorm

2019ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
Removes the re-centering (mean-subtraction) step of layer normalization — RMSNorm rescales activations using only the root-mean-square statistic, keeping re-scaling invariance while cutting LayerNorm's computational overhead.
Category
Architecture
Abstraction level
Building block
Operation level
LayerArchitecture block
Use cases
Pre-normalization in TransformersLarge language models (LLMs)QK-Norm (normalizing queries and keys in attention)Recurrent neural networks (RNNs)Vision and multimodal models

How it works

For an input vector a with n elements, RMSNorm computes the statistic RMS(a) = sqrt((1/n) · Σ aᵢ²) and then normalizes and rescales each component: āᵢ = (aᵢ / RMS(a)) · gᵢ, where g is a learnable gain vector of dimension n. Unlike LayerNorm, there is no mean subtraction and (in the basic version) no bias/offset term. A small constant ε is added to the denominator for numerical stability. In the partial variant (pRMSNorm), the RMS statistic is estimated from only the first p% of the vector's components, further lowering cost without breaking the invariance properties.

Problem solved

LayerNorm requires two passes over the activation vector — computing the mean (re-centering) and the variance (re-scaling) — which adds computational overhead that is especially costly in recurrent networks and deep Transformers. RMSNorm removes the re-centering step, reducing the number of operations and the running time while preserving training stability.

Components

RMS statisticScale measure of the activation vector used for normalization

The root mean square of the vector components: RMS(a) = sqrt((1/n)·Σ aᵢ²). It replaces LayerNorm's standard deviation and requires no mean computation.

Gain vector (g)Learnable, per-dimension rescaling of the normalized activations

A learnable parameter vector of dimension n, multiplied element-wise with the normalized activations. It provides adaptive re-scaling and implicit learning-rate adaptation.

Partial estimation (pRMSNorm)Optional approximation of the RMS statistic from p% of inputs

A variant where RMS is estimated from only the first p% of the vector components, further reducing compute cost without breaking the invariance properties.

Official

Implementation

Implementation pitfalls
Computing RMS in low precisionHigh

Summing squares in fp16/bf16 can lead to overflow or loss of precision.

Fix:Upcast activations to fp32 while computing the RMS statistic, then cast the result back.
Placement of epsilonMedium

Adding ε inside vs. outside the square root yields different numerical behavior and differs across implementations.

Fix:Follow the convention of the chosen reference implementation and stay consistent when porting weights.
Assuming mean subtraction or a bias termMedium

Porting LayerNorm code and leaving in mean subtraction or a bias term changes RMSNorm's semantics.

Fix:Remove the mean-computation step and (in the basic version) the bias parameter.

Evolution

Original paper · 2019 · NeurIPS 2019 · Biao Zhang
Root Mean Square Layer Normalization
Biao Zhang, Rico Sennrich
2016
Layer Normalization introduced

Ba, Kiros, and Hinton introduce LayerNorm — the predecessor that RMSNorm simplifies.

2019
RMSNorm published (NeurIPS 2019)
Inflection point

Zhang and Sennrich show that dropping re-centering matches LayerNorm while reducing running time by 7–64%.

2023
Adoption in LLaMA and other LLMs
Inflection point

LLaMA uses RMSNorm in a pre-normalization setup, after which the technique becomes a standard in subsequent model families (Gemma, Mistral, Qwen).

Hyperparameters (configurable axes)

Normalized dimension (n)High

Size of the activation vector being normalized; typically equal to the model hidden dimension.

Epsilon (ε)Medium

A small constant added for numerical stability (to avoid division by zero).

Learnable gain (g)High

Whether a learnable gain vector is used; in practice almost always enabled.

Partial ratio (p)Low

Fraction of inputs used to estimate RMS in the pRMSNorm variant.

Computational complexity

Time complexity: O(n · d). Space complexity: O(d).

Compute bottleneck

Sum-of-squares reduction over the hidden dimension

The operation is memory-bandwidth bound: an element-wise squaring and a reduction over dimension d, not a matrix multiplication. It does not exploit tensor cores.

Execution paradigm

Primary mode
Dense

A dense operation applied to all activations with no conditional routing.

Activation pattern
All paths active
Routing mechanism

Parallelism

Parallelism level
Fully parallel

Each token's normalization is independent of the others, so it is fully parallel across tokens; within a token there is a reduction over the hidden dimension.

Scope
TrainingInferenceAcross tokens

Hardware requirements

Good fit

An element-wise-plus-reduction operation, memory-bandwidth bound; runs efficiently on any accelerator and on CPU.

Possible

Does not use tensor cores (no matrix multiply); usually fused with neighboring ops for efficiency.