Google DeepMind SL2T moves sign-language AI toward a split mobile runtime

Google DeepMind introduced SL2T, a sign-language-to-text model starting with ASL-to-English on Pixel, with related social posts explaining benchmarks, privacy, and Deaf-community input.

Retrieval answer

Google DeepMind announced SL2T, a sign-language-to-text model reaching Android/Pixel product surfaces, starting with ASL-to-English on Pixel 11. The cluster includes a primary blog and official social thread with model claims, benchmark/privacy details, and Deaf-community input. This is a material accessibility AI release with concrete product consequences and no prior coverage match.

New Runtime synthesiseditorial-diagram
A privacy-aware flow from camera input through on-device pose landmarks to server translation and streaming text in Gboard and Live Transcribe.
New Runtime synthesis: SL2T separates on-device landmark extraction from server-side sign-language translation.New Runtime synthesisOriginal source ->

Field note

Google DeepMind introduced SL2T, a sign-language-to-text model starting with ASL-to-English on Pixel, with related social posts explaining benchmarks, privacy, and Deaf-community input.

Why it matters

Google DeepMind announced SL2T, a sign-language-to-text model reaching Android/Pixel product surfaces, starting with ASL-to-English on Pixel 11. The cluster includes a primary blog and official social thread with model claims, benchmark/privacy details, and Deaf-community input. This is a material accessibility AI release with concrete product consequences and no prior coverage match.

New Runtime view

SL2T is a split accessibility runtime: local feature extraction narrows the media boundary before a frontier model performs language translation. The consequence is not only better benchmark translation, but a product architecture for camera-based assistants where raw human motion should not automatically become cloud media.

Mechanism: On-device MediaPipe Holistic extracts body, face, and hand landmarks; Gemini 2.5 Pro translates the landmark stream into text.

Architectural boundary: Primary perception and privacy-sensitive video handling sit on Android/Pixel while semantic translation is server-side.

Measured consequence: DeepMind reports BLEU improvement from 15.0 to 21.4 versus the prior best on YouTube-ASL/How2Sign-style datasets.

What remains open

  • Research preview, not a certified interpreter replacement.
  • Signer diversity, dialect coverage, latency, consent UX, and accessibility liability remain open.
  • Landmarks reduce raw-video exposure but can still encode sensitive expression patterns.

Sources

  • <https://deepmind.google/blog/putting-sign-language-ai-into-users-hands>
  • <https://x.com/GoogleDeepMind/status/1955645677943959647>

Recommendation

Google DeepMind introduced SL2T, a sign-language-to-text model starting with ASL-to-English on Pixel, with related social posts explaining benchmarks, privacy, and Deaf-community input.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01topicAi - New RuntimeExplore the ai topic hub.
  2. 02topicModels - New RuntimeExplore the models topic hub.
  3. 03topicEdge Ai - New RuntimeExplore the edge-ai topic hub.
  4. 04archiveField NotesOpen the latest editorial analysis.
  5. 05source ledgerSource LedgerInspect the public source evidence graph.

These links are also published in this page's JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate...

Open the JSON contract