Robots Atlas>ROBOTS ATLAS
Sign-Language-to-Text (SL2T)

Sign-Language-to-Text (SL2T)

1.0 · Family: SL2T
Google DeepMind model translating American Sign Language (ASL) into English text, deployed in Gboard and Live Transcribe.
✓ Active✓ Public accessMultimodalSpecialized AI📁 SL2T
Release date
12 August 2026
Access:HostedDeployment:☁ Cloud📱 On-device

Overview

Sign-Language-to-Text (SL2T) is a machine-translation model developed by the Google DeepMind Sign Language Team. Its first version, SL2T 1.0, was unveiled on 12 August 2026 and translates American Sign Language (ASL) into English text.

The model is a transformer-based network that conditions on a sequence representing a signer's movements and generates a text translation autoregressively. Its input is not a raw camera feed but a sequence of 2D landmark coordinates for 130 key points of the face, hands and body, extracted frame-by-frame using MediaPipe Holistic. It was trained on more than 100,000 hours of data across more than 50 sign languages, with roughly a quarter of the data in ASL.

Deployment and privacy

Video is processed on the user's device and then immediately discarded — only the abstract landmark representation is sent to Google's servers. The model runs in a hybrid mode: on-device pose detection and server-side translation. SL2T 1.0 is initially available on Pixel 11 phones in Gboard (text dictation) and Live Transcribe (conversation responses) and is currently free.

Governance (AISLAC)

Deployment is guided by the AI Sign Language Advisory Committee (AISLAC), an advisory body convened by Google that brings together Deaf advocacy organizations, academic and cultural institutions and interpreters (including the National Association of the Deaf, RIT/NTID, DPAN and the World Federation of the Deaf). The Google DeepMind team and AISLAC co-authored the "Joint Impact Report for SL2T 1.0", detailing the model's capabilities and limitations. High-stakes uses (medical, legal, academic, and multi-party conversations) are explicitly out of scope for this release.

Classification
MultimodalSpecialized AI
Family: SL2T
Access & deployment
Hosted
CloudOn-device
Weights: Closed
Key parameters
📥 Input: video

Technical specification

Modalities
⬇ Input
video
⬆ Output
text

Capabilities and applications

Native model capabilities
Live Translation
Real-time speech translation between multiple languages without interrupting the audio stream.
Category: speech
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding

Benchmark results

4 benchmarks
FLEURS-ASL (sd-test)
BLEURT · zero-shot
70 (BLEU 25)
📅 12 Aug 2026📄 AISLAC Joint Impact Report for SL2T 1.0
FLEURS-ASL (1-handed)
BLEURT · zero-shot
74 (BLEU 32)
📅 12 Aug 2026📄 AISLAC Joint Impact Report for SL2T 1.0
PRESTO-ASL
BLEURT · zero-shot
85 (BLEU 67)
📅 12 Aug 2026📄 AISLAC Joint Impact Report for SL2T 1.0
FSboard
accuracy (exact match)
64%
📅 12 Aug 2026📄 AISLAC Joint Impact Report for SL2T 1.0

Technical architecture

Deployment and security

🔒 Security / Enterprise
✓ Verified enterprise information
Updated: 24 Aug 2026↗ Security documentation