KL Divergence
How it works
For discrete distributions P and Q, KL divergence is defined as D_KL(P || Q) = sum_x P(x) log( P(x) / Q(x) ), and for continuous ones as the integral of the analogous expression. It is non-negative (D_KL >= 0) and equals zero if and only if P = Q (Gibbs' inequality). It is not symmetric, D_KL(P || Q) != D_KL(Q || P), and does not satisfy the triangle inequality, so it is not a metric. It relates to cross-entropy: H(P, Q) = H(P) + D_KL(P || Q), so minimizing cross-entropy over Q's parameters is equivalent to minimizing KL. In AI practice it is computed over softmax outputs: in distillation, KL between a teacher's softened distribution and the student is minimized; in RLHF (PPO) a KL penalty against a reference policy is added to limit model drift; in variational inference, KL between an approximating distribution and the posterior is minimized.
Problem solved
Training and evaluating probabilistic models require a way to measure how far a predicted distribution deviates from a target or reference distribution. KL divergence provides such an information-theoretic measure: it lets you define loss functions (minimizing KL against the data is equivalent to maximizing likelihood), transfer knowledge from a teacher model to a student (distillation), approximate distributions in variational inference, and keep a fine-tuned model close to a reference policy in RLHF.
Key mechanisms
Strengths & limitations
Components
The reference distribution (e.g. the data, the teacher's distribution, the target distribution) against which divergence is measured.
The model distribution (e.g. student predictions, the fine-tuned policy, the variational distribution) whose distance from P is assessed.
The expression log(P(x)/Q(x)) weighted by P(x); its expectation under P yields the KL divergence.
Implementation
D_KL(P||Q) differs from D_KL(Q||P): forward KL enforces mode-covering, reverse KL is mode-seeking; the wrong direction yields undesired behavior.
When Q(x)=0 but P(x)>0, the log(P/Q) term goes to infinity, causing numerical instability.
Evolution
The paper On Information and Sufficiency introduces the distribution-divergence measure later called KL divergence.
Variational autoencoders (Kingma, Welling) use a KL term to regularize the latent space against a prior.
Knowledge distillation (Hinton, 2015) and later RLHF use KL as a knowledge-transfer objective and a penalty keeping the policy close to a reference.
Hyperparameters (configurable axes)
Forward KL D(P||Q) enforces mode-covering, reverse KL D(Q||P) is mode-seeking.
A factor softening softmax distributions before computing KL in knowledge distillation.
Weight of the KL penalty against the reference policy in RLHF/PPO.
Computational complexity
Time complexity: O(K) dla K kategorii rozkładu dyskretnego. Space complexity: O(K) na rozkład o K kategoriach.
Knowledge distillation using KL on softened distributions (Hinton 2015) lets smaller models approach the quality of larger teachers. In RLHF/PPO the KL penalty coefficient (e.g. beta) controls the tradeoff between reward fit and closeness to the reference policy - too low causes reward hacking and drift, too high stalls learning.
Compute bottleneck
KL cost grows with the number of distribution categories (e.g. vocabulary size in a language model); otherwise it is a cheap element-wise operation.
Execution paradigm
All distribution components enter the sum; there is no selective activation.
KL is a mathematical operation on distributions, with no routing or conditional execution.
Parallelism
The element-wise operation and reduction are fully parallel over probability vectors/tensors.
Hardware requirements
KL is an element-wise operation on distributions (log, multiply, reduction), efficiently realized on any accelerator supporting tensor ops.