The robot no longer repeats the same motion thousands of times. It looks, reads your words, and decides on its own how to grab the cup.
The industrial robot up to now is blind repetition. An engineer records a trajectory - the arm turns 37 degrees, grips with an exact force - and the machine repeats it thousands of times in the same spot, down to the millimeter. If the part moves two centimeters to the side, the robot doesn't see it. It has no eyes, only memory for a single movement. VLA flips the logic: instead of memorizing a motion, the model looks at the scene in the moment, understands what the person wants, and decides on its own how to do it.
Inside, VLA is three different things fused into one. In goes an image from the camera - the raw pixels of the table, the cup, the hand. In goes a sentence in natural language - the task in words. And out comes a stream of numbers: how many degrees to turn each joint, how tight the gripper should squeeze, when to stop. The model is trained on a huge amount of recordings - people and robots performing thousands of small everyday tasks - until it learns the connection between what it sees, what it hears as an instruction, and the movement that should follow.
The first model that proved this even works outside ideal lab conditions was RT-2 by Google DeepMind (Google's artificial intelligence research unit) in 2023. After it came OpenVLA - an open model any researcher can download and tinker with - and pi0 by Physical Intelligence, a startup of former Google and Tesla (the electric car maker) engineers, focused exactly on robots that do housework: fold clothes, load the dishwasher.
The difference that actually matters is generalization. An old robot trained to lift a red cup recognizes only that - a blue cup is a different object to it, a different task, requiring a new program. A VLA model has seen so many variations of 'grabbing an object' during training that it transfers what it learned even to something it has literally never encountered. It hasn't memorized a movement. It has understood the task at the level of language and translates it into action on its own, fresh every time.
The fragile spot is speed and reliability. These models still run slower than a learned industrial program and make mistakes exactly where the old robot never does - because it doesn't think at all, it only repeats. VLA thinks every time - and thinking takes time and sometimes comes out wrong. So for now these systems live mostly in research centers and demos, not on the assembly line.
Here's the trick
For twenty years I've watched machines do exactly one thing, brilliantly well, and nothing else. A printing press cuts exactly this shape. A CNC mill follows exactly this file. Every new task means a new file, a new setup, a new person who has to sit down and explain to the machine exactly what to do, move by move. VLA is the first time I've seen a machine you simply tell what you want - and it closes the distance between the word and the action on its own.
I'm not naive about it - I've watched enough demos trained to look perfect right in front of a camera to know how hard this is outside a controlled room. But the direction is right, and it isn't specific to robots with arms. Every system I build - an agent that reads a screen and clicks, an automation that reads an email and decides what to do with it - is the same principle: vision plus language, straight into action, no human in the middle for every step. Robotics just shows it most literally, because here the mistake hits the floor with a crash.