Timeline

Google puts real-time sign-language translation into a mainstream phone

SL2T, trained on over 100,000 hours of video across 50-plus sign languages, ships in Gboard and Live Transcribe on the Pixel 11 from around 20 August.

  • Models & capabilities
  • Culture & impact
  • Minor

Google DeepMind said it was building sign-language-to-text translation into Gboard and Live Transcribe on the Pixel 11, starting around 20 August, with support for more devices to follow. The underlying model, which Google calls SL2T, converts hand and body movement into written text for dictation and captioning, and the company described the rollout as bringing sign-language AI “out of the lab and into consumer products for the first time” — a claim about first mainstream deployment, not about the underlying research.

Google said SL2T was trained on more than 100,000 hours of video spanning over 50 sign languages, about a quarter of it American Sign Language. Rather than processing a raw camera feed, the system works from pose-landmark locations — tracked points on the body and hands — which Google said was a deliberate privacy choice, since it avoids sending video of a user’s surroundings off-device. On the FLEURS-ASL benchmark, Google reported a zero-shot score of 70 BLEURT, which it described as well above any previously reported score on that test; this is a self-reported figure and has not been independently verified.

Sign-language recognition has lagged well behind spoken-language translation in mainstream consumer AI, in part because sign languages are visually complex and data is comparatively scarce. Putting a translation model directly into a widely used keyboard and transcription app, rather than a standalone accessibility tool, is the more significant part of the release: it puts the feature in front of users who were not already seeking out specialised software.