Robots Atlas>ROBOTS ATLAS
Evaluation

FLEURS-ASL

2024ActivePublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
The first extension of the massively multilingual FLEURS/FLORES benchmark to a sign language (American Sign Language, ASL), enabling comparable evaluation of sign-language video translation into text and speech across all languages supported by FLORES (200) and FLEURS (102).
Category
Evaluation
Abstraction level
System
Operation level
Evaluation (runtime)Data
Use cases
Evaluation of ASL-to-text (sign-to-text) translationBenchmarking sign-language understanding of multimodal modelsComparable multilingual evaluation of sign-language translationAssessing ASL caption alignment, retrieval and receptive comprehensionSign-language translation research

How it works

Five Certified Deaf Interpreters (CDIs) translate sentences from the FLORES/FLEURS corpus into ASL on video, so each recording has ready-made text and speech references in many languages. For sentence-level scoring, two metrics are recommended: BLEU (sacreBLEU with "intl" tokenization) and BLEURT (the BLEURT-20 model); at the discourse level only BLEU is used (BLEURT is not intended for long outputs), and for timed translation, Timed BLEU (and by extension Timed BLEURT). The baseline model follows the YouTube-ASL approach: a T5v1.1-Base encoder-decoder into which 85 3D MediaPipe Holistic skeleton landmarks are linearly projected, with a 34-second context window and timestamp and previous-text tokens. The benchmark covers sentence- and discourse-level translation, caption alignment, retrieval, and multiple-choice receptive comprehension tasks.

Problem solved

Sign language translation has historically been peripheral to mainstream machine translation research, and there was no standardized, multilingual benchmark that included a sign language. FLEURS-ASL provides a comparable point of reference to rigorously measure the ability of models (including multimodal models) to understand and translate ASL.

Components

CDI video corpusSource of ASL evaluation data

1,749 sentences across 495 discourses (7.49 h of video) recorded by 5 Certified Deaf Interpreters (CDIs).

FLORES/FLEURS parallel corpusTarget translation references

Sentences come from FLORES/FLEURS, so ASL translations are aligned with text in 200 languages and speech in 102 languages.

Evaluation task suiteBenchmark task definitions

Sentence- and discourse-level translation, caption alignment, retrieval, and multiple-choice receptive comprehension.

Evaluation metricsScoring protocol

BLEU (sacreBLEU, "intl" tokenization) and BLEURT (BLEURT-20) at sentence level; BLEU only at discourse level; Timed BLEU/BLEURT for timed translation.

Baseline model (T5 + MediaPipe Holistic)Benchmark baseline

T5v1.1-Base encoder-decoder with linearly projected 85 3D MediaPipe Holistic landmarks, a 34-second context window, and timestamp and previous-text tokens; trained on YouTube-ASL.

Implementation

Implementation pitfalls
BLEURT unsuitable for long outputsMedium

The author notes BLEURT is not intended for long outputs, so at the discourse level only BLEU should be reported.

Fix:At the discourse level use BLEU only; reserve BLEURT (BLEURT-20) for sentence-level scoring.

Evolution

Original paper · 2024 · arXiv 2024 · Garrett Tanzer
FLEURS-ASL: Including American Sign Language in Massively Multilingual Multitask Evaluation
Garrett Tanzer
2021
FLORES-101 — multilingual text translation benchmark

Multilingual machine translation benchmark (101 languages) providing the text reference basis.

2022
FLEURS — speech for 102 languages

Speech extension of FLORES: Few-shot Learning Evaluation of Universal Representations of Speech (Google).

2023
YouTube-ASL — ASL–English corpus

Large-scale, open-domain ASL–English corpus (~1000 h), the basis for the FLEURS-ASL baseline model.

2024
FLEURS-ASL — first sign language in FLEURS/FLORES
Inflection point

Introduces ASL into massively multilingual evaluation; reveals that frontier multimodal models do not understand ASL.