we_are_coded.by CODE · The world, decoded
БГ
Concept

VLA: from pixels to motion

The BasicsUpdated on 16 August 2026we are coded

The robot no longer repeats the same motion thousands of times. It looks, reads your words, and decides on its own how to grab the cup.

Checked on16 August 2026
In short: VLA (vision-language-action) is a model that looks through a camera, reads a task in plain human language, and immediately produces a robot's movement. No line of code in between. You say "pick up the blue cup, put it next to the plate" and the arm moves - without anyone having programmed that exact motion beforehand.

The industrial robot up to now is blind repetition. An engineer records a trajectory - the arm turns 37 degrees, grips with an exact force - and the machine repeats it thousands of times in the same spot, down to the millimeter. If the part moves two centimeters to the side, the robot doesn't see it. It has no eyes, only memory for a single movement. VLA flips the logic: instead of memorizing a motion, the model looks at the scene in the moment, understands what the person wants, and decides on its own how to do it.

Inside, VLA is three different things fused into one. In goes an image from the camera - the raw pixels of the table, the cup, the hand. In goes a sentence in natural language - the task in words. And out comes a stream of numbers: how many degrees to turn each joint, how tight the gripper should squeeze, when to stop. The model is trained on a huge amount of recordings - people and robots performing thousands of small everyday tasks - until it learns the connection between what it sees, what it hears as an instruction, and the movement that should follow.

The first model that proved this even works outside ideal lab conditions was RT-2 by Google DeepMind (Google's artificial intelligence research unit) in 2023. After it came OpenVLA - an open model any researcher can download and tinker with - and pi0 by Physical Intelligence, a startup of former Google and Tesla (the electric car maker) engineers, focused exactly on robots that do housework: fold clothes, load the dishwasher.

The difference that actually matters is generalization. An old robot trained to lift a red cup recognizes only that - a blue cup is a different object to it, a different task, requiring a new program. A VLA model has seen so many variations of 'grabbing an object' during training that it transfers what it learned even to something it has literally never encountered. It hasn't memorized a movement. It has understood the task at the level of language and translates it into action on its own, fresh every time.

The fragile spot is speed and reliability. These models still run slower than a learned industrial program and make mistakes exactly where the old robot never does - because it doesn't think at all, it only repeats. VLA thinks every time - and thinking takes time and sometimes comes out wrong. So for now these systems live mostly in research centers and demos, not on the assembly line.

The eye is wired directly to the hand - between them there's no line of code waiting to be written for every new movement.

Here's the trick

For twenty years I've watched machines do exactly one thing, brilliantly well, and nothing else. A printing press cuts exactly this shape. A CNC mill follows exactly this file. Every new task means a new file, a new setup, a new person who has to sit down and explain to the machine exactly what to do, move by move. VLA is the first time I've seen a machine you simply tell what you want - and it closes the distance between the word and the action on its own.

I'm not naive about it - I've watched enough demos trained to look perfect right in front of a camera to know how hard this is outside a controlled room. But the direction is right, and it isn't specific to robots with arms. Every system I build - an agent that reads a screen and clicks, an automation that reads an email and decides what to do with it - is the same principle: vision plus language, straight into action, no human in the middle for every step. Robotics just shows it most literally, because here the mistake hits the floor with a crash.

The visual is generated code art. No third-party images.
Official primary sources
→Google DeepMind: RT-2 - Vision-Language-Action models (July 2023)