Robots Atlas>ROBOTS ATLAS
Data

Gloss

2018ActivePublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
Introduced an intermediate, symbolic representation of sign language — a word-for-word transcription using spoken-language words in capital letters — that lets recognition and translation systems treat sign language as a sequence-to-sequence problem.
Category
Data
Abstraction level
Building block
Operation level
DataTrainingInference
Use cases
Sign language translation (SLT)Continuous sign language recognition (CSLR)Sign language corpus annotationIntermediate supervision target (CTC) in model trainingSign language linguistics research

How it works

An annotator writes each manual sign as a spoken-language word in ALL CAPS in signing order; non-manual features and prosody are noted as superscripts with scope in brackets, fingerspelling with hyphens, and lexicalized forms with #. In the ML pipeline: (1) sign2gloss2text — a visual network recognizes a gloss sequence from video (usually trained with a CTC loss since there is no frame-level alignment) and a second network translates the glosses into a sentence; (2) in joint models (Sign Language Transformers) the gloss acts as an auxiliary CTC supervision target alongside the main translation objective. Because gloss order differs from spoken word order (non-monotonicity), the gloss-to-text stage performs full sequence-to-sequence translation.

Problem solved

Sign languages have no widely adopted written form, and continuous signing is hard to align to spoken text. A gloss provides a discrete, tokenized mid-level representation that enables supervised recognition and translation and quantitative evaluation of systems.

Components

Manual gloss token (ALL CAPS)Basic representational unit: a spoken-language word denoting a manual sign

A manual sign written as a capitalized spoken-language word, placed in signing order.

Non-manual / prosody markersEncoding facial expression, head movement, negation and questions and their scope

Usually written as superscripts with scope in brackets, e.g. [I LIKE]^negative.

Fingerspelling notationNotation of finger-spelled letters and lexicalized forms

Pure fingerspelling marked with hyphens (W-I-K-I); lexicalized forms marked with a hash (#JOB).

Spatial/indexing notationMarking spatial reference and agreement (indexing, IX)

Partial encoding of spatial grammar; inherently incomplete, which is a source of information loss.

Implementation

Implementation pitfalls
Information loss (gloss is not a full language)High

Gloss flattens simultaneity, spatial grammar and non-manual features, creating an information bottleneck that limits translation quality.

Fix:Use gloss-free approaches or richer, multi-layer non-manual annotation.
High annotation costHigh

Gloss annotation requires trained linguists or native signers; datasets remain small (on the order of 10–20k pairs).

Fix:Transfer learning, weak supervision, YouTube-scale data and gloss-free learning.
No universal glossing standardMedium

Conventions differ across corpora and sign languages, hurting reproducibility and cross-dataset transfer.

Fix:Document the glossing convention and map vocabularies across corpora.
Non-monotonicity vs. spoken languageMedium

Gloss order differs from spoken word order, so gloss-to-text alignment is non-trivial.

Fix:Implement the gloss-to-text stage as full sequence-to-sequence translation.

Evolution

Original paper · 2018 · CVPR 2018 · Necati Cihan Camgöz
Neural Sign Language Translation
Necati Cihan Camgöz, Simon Hadfield, Oscar Koller, Hermann Ney, Richard Bowden
2014
RWTH-PHOENIX-Weather 2014 gloss-annotated corpus

German TV weather forecasts annotated with glosses — a standard benchmark for continuous sign language recognition.

2018
Neural Sign Language Translation introduces sign2gloss2text and PHOENIX-2014T
Inflection point

Camgöz et al. formalize gloss as a mid-level representation in neural sign language translation.

2020
Sign Language Transformers: gloss as intermediate CTC supervision

Joint end-to-end recognition and translation; gloss as an auxiliary CTC target drastically improves translation quality.

2022
Multi-modality transfer baseline (sign2gloss2text) reaches strong results

Chen et al. connect a sign-to-gloss and a gloss-to-text network with a visual-language mapper, boosting the gloss pipeline via transfer learning.

2023
GFSLT-VLP: gloss-free translation and critique of the gloss bottleneck
Inflection point

Zhou et al. point to the "information bottleneck" of gloss representation and data scarcity, proposing gloss-free translation.

2024
FLEURS-ASL: gloss-free video benchmark for ASL

Extends FLORES/FLEURS with American Sign Language as video; shows frontier models have virtually no understanding of ASL.

Hyperparameters (configurable axes)

Glossing conventionHigh

No universal standard — conventions differ across corpora and sign languages (ASL, DGS, BSL).

Gloss vocabulary sizeMedium

The number of unique glosses in a corpus affects recognition difficulty and coverage.

Non-manual annotationMedium

Whether and how finely facial expression, prosody and spatial grammar are noted alongside manual glosses.