The model watches body points, not the video, and turns American Sign Language into text on the device itself. The frame is discarded immediately.
- On 12 August DeepMind announced the SL2T model inside Gboard and Live Transcribe on Pixel 11.
- It was trained on over 100,000 hours of material across more than 50 sign languages.
- American Sign Language to English goes first; others follow.
What stopped me first was the design, not the news itself. The model does not watch your video, it tracks where the joints of your hands and body are.
The small technical move here is also the privacy decision. Once coordinates go into processing instead of a face, the whole argument about what happens to the frame dissolves. Google says the video itself is discarded as soon as the points are read.
It works on one phone, for one language, in two apps. That is all for now.