Evaluation
Perplexity
1977ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key
innovation
A measure of language-model quality: the exponential of the average negative log-likelihood — intuitively the "effective number of equally likely choices" per prediction step.
Category
Evaluation
Abstraction level
Primitive
Operation level
Evaluation (runtime)Training
Use cases
Evaluating language models during pretrainingComparing architectures on the same corpusMonitoring training convergenceAssessing compression/distribution fit
How it works
For a test corpus, one computes the log-probability of each token under the model, averages the negative log-probabilities (cross-entropy) and exponentiates. Because it depends on tokenisation and vocabulary, fair comparisons require identical data and tokenisation; lower perplexity means better fit.
Problem solved
One needs to measure how well a model predicts text, independent of a specific task. Perplexity gives a single, comparable number reflecting language-modelling quality.
Components
Cross-entropyMetric basis
Average negative log-likelihood of the data under the model.
ExponentiationInterpretation
Turns cross-entropy into an "effective number of choices".
Evolution
1977
Perplexity introduced in speech recognition (IBM, Jelinek et al.)
Inflection point2018
Standard pretraining metric for modern LLMs