we_are_coded.by CODE · The world, decoded
БГ
Google DeepMind

Google released EmbeddingGemma 2: an open 740-million-parameter model for searching text, code, images, video and audio on the device itself

Google · event date: 6 October 2026Frontier

One shared “shelf” for five kinds of content, which can run entirely on the device. By Google’s figures the full multimodal version, compressed, needs about 567 MB of memory on a Pixel 11 Pro. The licence is Apache 2.0, and the weights are on Hugging Face and Kaggle.

In short
  • EmbeddingGemma 2 puts text, code, images, video and audio in one shared space. 740 million parameters, Gemma 4 architecture, Apache 2.0 licence.
  • By Google’s figures it needs about 567 MB of active memory for the compressed full version on a Pixel 11 Pro, and it also works offline.
  • The weights are on Hugging Face and Kaggle. The first version, text only, has over 20 million downloads, by Google’s figures.
Checked on7 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

You record a voice memo and use it to find the video clip you are talking about. The example is Google’s own, and it beats any table.

Here is what an embedding is, in one picture. Imagine a library where things are not shelved alphabetically but by meaning: the note about the sea, the photo of the sea and the song about the sea sit on one shelf. The model shelves everything, and search asks what sits closest to your question.

The facts: on 6 October 2026 Google DeepMind (authors Sahil Dua and Henrique Schechter Vera) released EmbeddingGemma 2 - an open, lightweight model for multimodal embeddings that maps text, code, images, video and audio into one shared space. Built on the Gemma 4 architecture, Apache 2.0 licence, 740 million parameters. The first version, from last year, is text only and by Google’s figures has over 20 million downloads. The context is 8K tokens: up to 5.5 minutes of audio, 29 images or 58 video frames. By Google’s figures, with quantization (compressing the weights) on a Pixel 11 Pro the full multimodal model needs about 567 MB of active memory. By Google’s figures the MTEB Code score is 78.68, against 68.76 for its predecessor. The weights are on Hugging Face and Kaggle, with Model Garden on the Gemini Enterprise Agent Platform coming soon.
You do not remember what the file is called. You remember what it is about.

The number that concerns you is 567 MB. That is what the compressed full version needs on one phone, by Google’s figures. At that size, searching your photos, recordings and clips may never have to leave the phone. Google presents it as a win for privacy and speed: the sorting happens on the spot and works without internet too.

Code is the second use. On the code benchmark Google reports a jump of almost 10 points over the first version and points to local codebase indexing and search for coding agents.

I haven't run it. The numbers are Google’s, and the memory figure comes from one phone model. The licence is Apache 2.0 and allows commercial use, so anyone building an app for photos or notes now has something to search them with, without sending them to the cloud.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Google - EmbeddingGemma 2: an open, lightweight multimodal embedding model, 06.10.2026→Hugging Face - google/embeddinggemma-2 (the weights and the Apache 2.0 licence)
Original: https://wearecoded.com/en/articles/google-embeddinggemma-2-740m-multimodalen.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news