we_are_coded.by CODE · The world, decoded
БГ
Concept

Embedding: when meaning gets an address

The BasicsUpdated on 7 October 2026we are coded

An embedding turns a word, a photo or a sound into a row of numbers - something like an address on a map of meaning. Things close in meaning get close addresses. That's why a search finds "cheap hotel by the sea" even when the text says "budget room on the beach".

Checked on7 October 2026
In short: an embedding (a vector representation) is a list of numbers a model uses to describe the meaning of a piece of content - a sentence, a photo, a sound. Those numbers work like coordinates: things with similar meaning get nearby coordinates, distant ones get distant coordinates. The list is called a vector. The model that makes it doesn't write answers, it only arranges. Search by meaning, RAG and finding a photo by its description all stand on it.

Picture a library where the books aren't shelved alphabetically but by meaning. The cookbooks sit in one corner, and the book on Italian cooking stands a hand's width from the one on pasta, even though their titles share no word. An embedding does exactly that with every sentence and every photo: it gives them a place in a library like that. The place is written down in numbers, hundreds of them (in EmbeddingGemma 2 up to 768), so nobody can draw it, but a machine can measure it.

Old search compares letters. You type "car" and it finds the texts where "car" appears. A text about an "automobile" stays out. Search with embeddings compares addresses. Your question also gets a place in the library, and the machine simply takes the books standing closest to it. That's why it finds the answer even when the words differ. And that's also why it sometimes finds something close but not the exact thing: closeness isn't the same as being right.

Where you meet it without seeing it

In a site search that understands what you're asking. In RAG, the setup where a chatbot first searches your documents and only then answers: it's the embedding that decides which passages are on topic. On your phone, when you type "dog on a beach" in the gallery and photos come up that nobody ever tagged. Anywhere a machine looks for the similar thing without an exact match.

The reason we're explaining it now is a small model. On October 6, 2026 Google released EmbeddingGemma 2: an open model with 740 million parameters, built on the Gemma 4 architecture, under the Apache 2.0 license, which also allows commercial use. It puts text, code, images, video and audio into one shared space, one library for everything. That way a voice memo can find a video clip, and a text question can find a spot in hours of audio recordings. It's built to run on the device itself, even offline. By Google's figures, the first EmbeddingGemma, text only, was downloaded more than 20 million times.

An embedding doesn't know what's true. It only knows what sits next to what.

That's why the small size is the news. A model that arranges meaning on your phone means your photos, notes and recordings can be searched by meaning without leaving the phone. For your data, that weighs more than any number in the benchmark tables. The test is the day you meet it in an app you already use.

The visual is generated code art. No third-party images.
Official primary sources
→Google: EmbeddingGemma 2 (Oct 6, 2026)