Robots Atlas>ROBOTS ATLAS
Training

KL Divergence

1951ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
An information-theoretic measure quantifying how much one probability distribution differs from a second (reference) one — interpreted as the average number of extra bits needed to encode data from distribution P using a code optimized for Q.
Category
Training
Abstraction level
Primitive
Operation level
TrainingPost-trainingEvaluation (runtime)
Use cases
Loss function and regularization in model trainingKnowledge distillation (KL between teacher and student)KL penalty against a reference policy in RLHF/PPOThe ELBO bound in variational inference (VAE)Measuring output-distribution divergence in model quantization and compression

How it works

For discrete distributions P and Q, KL divergence is defined as D_KL(P || Q) = sum_x P(x) log( P(x) / Q(x) ), and for continuous ones as the integral of the analogous expression. It is non-negative (D_KL >= 0) and equals zero if and only if P = Q (Gibbs' inequality). It is not symmetric, D_KL(P || Q) != D_KL(Q || P), and does not satisfy the triangle inequality, so it is not a metric. It relates to cross-entropy: H(P, Q) = H(P) + D_KL(P || Q), so minimizing cross-entropy over Q's parameters is equivalent to minimizing KL. In AI practice it is computed over softmax outputs: in distillation, KL between a teacher's softened distribution and the student is minimized; in RLHF (PPO) a KL penalty against a reference policy is added to limit model drift; in variational inference, KL between an approximating distribution and the posterior is minimized.

Problem solved

Training and evaluating probabilistic models require a way to measure how far a predicted distribution deviates from a target or reference distribution. KL divergence provides such an information-theoretic measure: it lets you define loss functions (minimizing KL against the data is equivalent to maximizing likelihood), transfer knowledge from a teacher model to a student (distillation), approximate distributions in variational inference, and keep a fine-tuned model close to a reference policy in RLHF.

Key mechanisms

Expectation of log(P/Q) under P
Relation to cross-entropy: H(P,Q) = H(P) + D_KL(P||Q)
Minimizing KL as a training / distillation objective
KL penalty against a reference policy (RLHF/PPO)
KL term in the evidence lower bound ELBO (VAE)

Strengths & limitations

Strengths
✓Grounded in information theory (bit interpretation)
✓Related to cross-entropy and maximum-likelihood estimation
✓Universal: loss functions, distillation, RLHF, VAE, quantization measurement
✓Differentiable, easy to optimize by gradient descent
✓Computationally cheap (element-wise operation on distributions)
Limitations
✗Asymmetric: D_KL(P||Q) != D_KL(Q||P) - requires choosing a direction
✗Not a metric (no triangle inequality)
✗Diverges to infinity when Q(x)=0 but P(x)>0
✗Numerically sensitive without log-softmax / smoothing
✗Forward vs reverse KL give different behaviors (mode-covering vs mode-seeking)

Components

Base distribution PWeights in the KL sum/integral

The reference distribution (e.g. the data, the teacher's distribution, the target distribution) against which divergence is measured.

Approximating distribution QThe optimization target that minimizes KL

The model distribution (e.g. student predictions, the fine-tuned policy, the variational distribution) whose distance from P is assessed.

Log-likelihood ratioThe computational core of the measure

The expression log(P(x)/Q(x)) weighted by P(x); its expectation under P yields the KL divergence.

Implementation

Implementation pitfalls
Asymmetry and choice of directionMedium

D_KL(P||Q) differs from D_KL(Q||P): forward KL enforces mode-covering, reverse KL is mode-seeking; the wrong direction yields undesired behavior.

Fix:Deliberately pick the direction to match the goal; consider symmetric alternatives (Jensen-Shannon divergence) when symmetry is needed.
Divergence to infinity at zero QHigh

When Q(x)=0 but P(x)>0, the log(P/Q) term goes to infinity, causing numerical instability.

Fix:Use smoothing, log-softmax and stable implementations; operate in log-probability space.

Evolution

Original paper · 1951 · Annals of Mathematical Statistics · Solomon Kullback
On Information and Sufficiency
Solomon Kullback, Richard Leibler
1951
Kullback and Leibler define the divergence
Inflection point

The paper On Information and Sufficiency introduces the distribution-divergence measure later called KL divergence.

2013
KL as a term in the VAE ELBO

Variational autoencoders (Kingma, Welling) use a KL term to regularize the latent space against a prior.

2017
Distillation and KL penalty in RLHF/PPO

Knowledge distillation (Hinton, 2015) and later RLHF use KL as a knowledge-transfer objective and a penalty keeping the policy close to a reference.

Hyperparameters (configurable axes)

Divergence directionHigh

Forward KL D(P||Q) enforces mode-covering, reverse KL D(Q||P) is mode-seeking.

forward (P||Q)Mode-covering, e.g. cross-entropy.
reverse (Q||P)Mode-seeking, e.g. variational inference.
Temperature (distillation)Medium

A factor softening softmax distributions before computing KL in knowledge distillation.

T=1No softening.
T=2-4Typical in distillation.
KL penalty coefficient (RLHF)High

Weight of the KL penalty against the reference policy in RLHF/PPO.

beta ~ 0.01-0.1Balance of reward vs drift.

Computational complexity

Computational characteristics
→Element-wise operation: log, multiply, reduction (sum)
→Cost linear in the number of categories (e.g. vocabulary size)
→Differentiable with respect to Q's parameters
→Stabilized by log-softmax and smoothing
→Efficiently computed on CPU and GPU

Time complexity: O(K) dla K kategorii rozkładu dyskretnego. Space complexity: O(K) na rozkład o K kategoriach.

Benchmark notes

Knowledge distillation using KL on softened distributions (Hinton 2015) lets smaller models approach the quality of larger teachers. In RLHF/PPO the KL penalty coefficient (e.g. beta) controls the tradeoff between reward fit and closeness to the reference policy - too low causes reward hacking and drift, too high stalls learning.

Compute bottleneck

Event-space size (vocabulary)

KL cost grows with the number of distribution categories (e.g. vocabulary size in a language model); otherwise it is a cheap element-wise operation.

Depends on
Liczba kategorii KStabilność numeryczna

Execution paradigm

Primary mode
Dense

All distribution components enter the sum; there is no selective activation.

Activation pattern
All paths active
Routing mechanism

KL is a mathematical operation on distributions, with no routing or conditional execution.

Parallelism

Parallelism level
Fully parallel

The element-wise operation and reduction are fully parallel over probability vectors/tensors.

Scope
TrainingInference

Hardware requirements

Primary

KL is an element-wise operation on distributions (log, multiply, reduction), efficiently realized on any accelerator supporting tensor ops.