Robots Atlas>ROBOTS ATLAS
Evaluation

Perplexity

1977ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
A measure of language-model quality: the exponential of the average negative log-likelihood — intuitively the "effective number of equally likely choices" per prediction step.
Category
Evaluation
Abstraction level
Primitive
Operation level
Evaluation (runtime)Training
Use cases
Evaluating language models during pretrainingComparing architectures on the same corpusMonitoring training convergenceAssessing compression/distribution fit

How it works

For a test corpus, one computes the log-probability of each token under the model, averages the negative log-probabilities (cross-entropy) and exponentiates. Because it depends on tokenisation and vocabulary, fair comparisons require identical data and tokenisation; lower perplexity means better fit.

Problem solved

One needs to measure how well a model predicts text, independent of a specific task. Perplexity gives a single, comparable number reflecting language-modelling quality.

Components

Cross-entropyMetric basis

Average negative log-likelihood of the data under the model.

ExponentiationInterpretation

Turns cross-entropy into an "effective number of choices".

Evolution

1977
Perplexity introduced in speech recognition (IBM, Jelinek et al.)
Inflection point
2018
Standard pretraining metric for modern LLMs