A language model predicts the next word. A world model predicts the next moment: where the cup will fall, how the robot's arm will turn, what will appear on screen after your click. This is where robotics, video and interfaces meet.
When you throw a ball to a child, they don't calculate parabolas. They just know where it will land, because they've watched things fall thousands of times. A world model builds the same feel from video. It watches enough to carry the scene forward on its own.
Three examples from one autumn
On September 22, 2026 Black Forest Labs, the makers of the FLUX image generator, released FLUX 3 Action: an open-weights model with 7 billion parameters that predicts a robot's future frames and actions at the same time. They call it a world action model themselves.
NVIDIA uses its Cosmos platform for robots and driverless cars: it simulates how the physical world behaves, so the system can be trained and tested without the real road. And on August 31, 2026 Runway showed Solaris, the first model in a family it calls Interface World Models. There the world is the interface itself: the app is drawn frame by frame while you use it.
Honest about the limits
Predicting the frame doesn't yet mean understanding the physics. A model can show a convincing video where the impossible looks normal. That's why robotics measures these models by how many tasks the robot completes. FLUX 3 Action, for example, solves just over 42% of the tasks in RoboLab-120, a simulation benchmark, by the company's numbers.