Robots Atlas>ROBOTS ATLAS
Evaluation

BLEURT

2020ActivePublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
Replaced n-gram overlap metrics (BLEU, ROUGE) with a learned BERT-based model that predicts human ratings and correlates with them far better, gaining robustness from pre-training on millions of synthetic sentence pairs.
Category
Evaluation
Abstraction level
Building block
Operation level
Evaluation (runtime)Inference
Use cases
Machine translation quality evaluationText generation evaluation (data-to-text, summarization, dialogue)Sign-language translation evaluation (FLEURS-ASL, Sign-Language-to-Text)Benchmarking and comparing NLG modelsAutomatic scoring of paraphrases and semantic equivalence

How it works

BLEURT is a regression model built on top of BERT and trained in two stages. (1) Pre-training on millions of synthetic sentence pairs produced by randomly perturbing Wikipedia sentences (e.g. token masking, permutations, backtranslation); these pairs are not human-rated — labels are generated by a collection of existing signals and models from the literature (e.g. BLEU, ROUGE, BERTScore, backtranslation likelihood, and textual entailment signals). This teaches the model general similarity assessment and makes it robust to distribution shift. (2) Fine-tuning on a small set of human ratings (WMT Metrics Shared Task, WebNLG). At inference the model takes a (candidate, reference) pair, encodes it with the BERT/RemBERT encoder, and a regression head over the sentence representation returns a single number — the predicted quality score (in practice roughly 0 for a random output and 1 for a perfect one).

Problem solved

Traditional metrics such as BLEU and ROUGE measure surface-level n-gram overlap between a model output and a reference, so they correlate poorly with human judgment and fail to capture paraphrases or semantic equivalence. BLEURT addresses this by learning to predict human ratings and to capture meaning-level similarity rather than literal word overlap, while remaining robust to out-of-distribution data.

Components

BERT / RemBERT encoderSemantic encoding of the sentence pair

A pre-trained Transformer encoder that encodes the (candidate, reference) pair into contextual representations. Monolingual checkpoints use BERT, while the multilingual BLEURT-20 uses RemBERT.

Official

Regression headQuality score prediction

A linear layer on top of the sentence representation (the [CLS] token) that maps the representation to a single number — the predicted quality score.

Synthetic-signal pre-trainingRobustness and generalization

The pre-training stage on millions of synthetic sentence pairs (Wikipedia perturbations) automatically labelled by existing metrics and models (BLEU, ROUGE, BERTScore, backtranslation, textual entailment). Provides robustness and generalization with few human ratings.

Official

Human-rating fine-tuningCalibration to human judgment

Fine-tuning on a small set of real human ratings (WMT Metrics Shared Task, WebNLG), calibrating predictions to actual quality judgments.

Official

Implementation

Implementation pitfalls
Scores are not strictly bounded to 0–1Medium

Raw BLEURT values are regression predictions and can fall outside the 0–1 range; ~0 and ~1 are only rough reference points (random vs perfect output).

Fix:Treat the score as relative; do not assume a hard range or interpret absolute values as percentages.
Scores are not comparable across checkpointsHigh

Different checkpoints (BERT-base, BERT-large, BLEURT-20, BLEURT-20-D12) produce values on different scales; comparing scores across checkpoints is invalid.

Fix:Always report the checkpoint used and compare scores only within the same checkpoint.
Dependence on reference and fine-tuning dataMedium

BLEURT requires a reference and is fine-tuned on human ratings (mostly WMT), so it can carry domain/language biases and perform worse out of distribution.

Fix:Validate the metric on your own domain/language; for reference-free needs consider reference-free metrics (e.g. COMET-QE).

Evolution

Original paper · 2020 · ACL 2020 · Thibault Sellam
BLEURT: Learning Robust Metrics for Text Generation
Thibault Sellam, Dipanjan Das, Ankur P. Parikh
2020
BLEURT published at ACL 2020
Inflection point

Sellam, Das and Parikh introduce a learned BERT-based metric, pre-trained on synthetic sentence pairs and fine-tuned on human ratings, with higher correlation to human judgment than BLEU.

2021
BLEURT-20: multilingual and distilled metrics (RemBERT)

The paper 'Learning Compact Metrics for MT' (Pu et al., EMNLP 2021) advances RemBERT-based multilingual metrics and distillation; it underpins the recommended BLEURT-20 checkpoint and its lighter variants (e.g. BLEURT-20-D12).

Hyperparameters (configurable axes)

Base checkpointCritical

Choice of the pre-trained encoder: BERT (monolingual) or RemBERT (multilingual, BLEURT-20). Determines language coverage and quality.

BERT (BLEURT base/large)
RemBERT (BLEURT-20)
Model sizeHigh

Number of parameters / layers. Larger checkpoints yield better correlation at the cost of speed; distilled variants (e.g. BLEURT-20-D12) lower cost.

BLEURT-20 (pełny)
BLEURT-20-D12 (destylowany, 12 warstw)
Maximum sequence lengthMedium

Token limit for the concatenated candidate + reference pair; longer inputs are truncated, which can affect scoring of very long texts.

Computational complexity

Time complexity: O(n² · d). Space complexity: O(n² + n · d).

Execution paradigm

Primary mode
Dense

BLEURT uses a dense Transformer encoder (BERT/RemBERT) — all parameters are active for every input; there is no conditional routing.

Activation pattern
All paths active

Parallelism

Parallelism level
Fully parallel

The encoder processes all tokens of the pair in parallel in a single pass; scoring many pairs batches easily.

Scope
InferenceTrainingAcross tokens

Hardware requirements

Primary

Scoring is a Transformer encoder forward pass; GPUs with tensor cores greatly accelerate matrix multiplications and batching of pairs.

Good fit

The reference implementation uses TensorFlow and maps well onto TPUs for training and large-scale scoring.

Possible

BLEURT runs on CPU but slower; useful for small numbers of pairs or when no accelerator is available.