Google DeepMind has unveiled a breakthrough in sign language AI with the introduction of SL2T, a massively multilingual sign-language-to-text translation model. For the first time, this technology is being integrated into consumer products, enabling Deaf and hard of hearing users to dictate by signing rather than typing.
The model powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11 devices, initially supporting American Sign Language (ASL) to English translation. Users can sign to search the web, draft messages, or interact with Gemini, while Live Transcribe allows signing responses in conversations. Early testers report that signing in ASL is faster and more natural than typing in English.
SL2T addresses two core challenges: sign languages are independent languages with distinct grammars, requiring true machine translation, and the model must accurately track simultaneous movements of hands, arms, torso, head, and face. The model is trained on over 100,000 hours of data across more than 50 sign languages, with about a quarter in ASL, and translates directly from pose landmarks to text, bypassing intermediate gloss annotations.
Privacy is a priority: SL2T processes only geometric coordinates of the signer's body, not raw video, and the original video is discarded immediately. The model achieves a zero-shot score of 70 BLEURT on the FLEURS-ASL benchmark, surpassing previous results. The team also addressed practical issues like latency, hallucination prevention, and fairness for left-handed and one-handed signing.
Developed in collaboration with the Deaf community, the project was guided by the AI Sign Language Advisory Committee (AISLAC) and includes a joint impact report. The feature is available at no additional cost on Pixel 11, with more devices and languages planned.