we_are_coded.by CODE · The world, decoded
БГ
Concept

Multimodal model

The BasicsUpdated on 16 August 2026we are coded

A model that reads not just words but also photos, sound, video - as if it has more than one sense.

Checked on16 August 2026
In short: an ordinary AI model understands only words, like reading a letter without seeing the photo taped to it. A multimodal model has more than one sense: it reads, looks at photos, listens to sound, follows video. That's why it can turn a sentence into a picture. And in robots, it turns what the camera sees into a movement of the arm.

Imagine describing over the phone to a friend what the dress you bought looks like. She listens, but sees nothing. Now imagine you also send her a photo. Suddenly she understands far more, with fewer words. That's exactly the difference between a model that only reads text and one that, besides text, also sees a picture. The second sense doesn't replace the first, it adds to it.

This is where the phrase text-to-image comes from, literally from words to a picture. You type a sentence like "a cat sitting on a windowsill, evening light" and the model draws exactly that. Not because it's an artist in the classic sense, but because it has linked the words we've shown it thousands of times with the photos that came alongside them. It has learned what looks like what.

With robots things get more tangible. A robot in a warehouse has a camera instead of eyes. The camera sends pixels, those same dots that make up every digital photo, to the model. The model looks at those pixels and decides: that box is to the left, the arm needs to turn that way and grab it. From pixels to action, without a person pressing a button at every step.

A multimodal model isn't smarter because it knows more words. It's smarter because it links senses together, the way we do.

Through the eye, not only through words

I'm a designer, so this topic is close to me from the inside. I've spent my whole life working with people who explain an idea in words, and I only see it once they show me a moodboard or a sketch. Text on its own has always been insufficient for me.

So when AI stopped only reading text and started looking at photos, that wasn't a technical detail to me. It was the moment the machine started thinking a little closer to how I think - through the eye, not only through words.

The visual is generated code art. No third-party images.
Official primary sources
→Google: Gemini - natively multimodal (6 December 2023)