One shared “shelf” for five kinds of content, which can run entirely on the device. By Google’s figures the full multimodal version, compressed, needs about 567 MB of memory on a Pixel 11 Pro. The licence is Apache 2.0, and the weights are on Hugging Face and Kaggle.
- EmbeddingGemma 2 puts text, code, images, video and audio in one shared space. 740 million parameters, Gemma 4 architecture, Apache 2.0 licence.
- By Google’s figures it needs about 567 MB of active memory for the compressed full version on a Pixel 11 Pro, and it also works offline.
- The weights are on Hugging Face and Kaggle. The first version, text only, has over 20 million downloads, by Google’s figures.
You record a voice memo and use it to find the video clip you are talking about. The example is Google’s own, and it beats any table.
Here is what an embedding is, in one picture. Imagine a library where things are not shelved alphabetically but by meaning: the note about the sea, the photo of the sea and the song about the sea sit on one shelf. The model shelves everything, and search asks what sits closest to your question.
The number that concerns you is 567 MB. That is what the compressed full version needs on one phone, by Google’s figures. At that size, searching your photos, recordings and clips may never have to leave the phone. Google presents it as a win for privacy and speed: the sorting happens on the spot and works without internet too.
Code is the second use. On the code benchmark Google reports a jump of almost 10 points over the first version and points to local codebase indexing and search for coding agents.
I haven't run it. The numbers are Google’s, and the memory figure comes from one phone model. The licence is Apache 2.0 and allows commercial use, so anyone building an app for photos or notes now has something to search them with, without sending them to the cloud.