A model that reads not just words but also photos, sound, video - as if it has more than one sense.
Imagine describing over the phone to a friend what the dress you bought looks like. She listens, but sees nothing. Now imagine you also send her a photo. Suddenly she understands far more, with fewer words. That's exactly the difference between a model that only reads text and one that, besides text, also sees a picture. The second sense doesn't replace the first, it adds to it.
This is where the phrase text-to-image comes from, literally from words to a picture. You type a sentence like "a cat sitting on a windowsill, evening light" and the model draws exactly that. Not because it's an artist in the classic sense, but because it has linked the words we've shown it thousands of times with the photos that came alongside them. It has learned what looks like what.
With robots things get more tangible. A robot in a warehouse has a camera instead of eyes. The camera sends pixels, those same dots that make up every digital photo, to the model. The model looks at those pixels and decides: that box is to the left, the arm needs to turn that way and grab it. From pixels to action, without a person pressing a button at every step.
Through the eye, not only through words
I'm a designer, so this topic is close to me from the inside. I've spent my whole life working with people who explain an idea in words, and I only see it once they show me a moodboard or a sketch. Text on its own has always been insufficient for me.
So when AI stopped only reading text and started looking at photos, that wasn't a technical detail to me. It was the moment the machine started thinking a little closer to how I think - through the eye, not only through words.