
Google’s artificial intelligence lab announced on the 12th local time the launch of a large-scale multilingual sign language-to-text (SL2T) translation model.
Thanks to this model, sign language AI has entered consumer products for the first time: the Gboard and Live Transcribe apps on Google Pixel 11 series smartphones support sign language-to-text input (initially limited to American Sign Language, with support for more sign languages and devices to be added later).
Approximately 70 million people with hearing impairments worldwide use more than 200 sign languages. The new feature enables them to enjoy a flexible “typing alternative” similar to voice input. Tests show that using American Sign Language is faster, more natural, and more enjoyable than typing in English.

DeepMind said that sign language is an independent natural language with its own unique grammar and vocabulary. Therefore, sign language translation requires genuine machine translation rather than converting signs into text sequentially. A model capable of performing SL2T tasks must be able to “understand” the synchronized movements of multiple parts of a sign language user’s body, requiring powerful computer vision capabilities.
Google DeepMind’s SL2T model was trained on 100,000 hours of data covering more than 50 sign languages. It converts sign language into positional information for a series of pose keypoints to protect user privacy, without sending the original video stream to servers. It also abandons the widely used “intermediate annotations” of the past and directly interprets spatial information.
