Google's new model tracks its own work through video and corrects itself on the fly. It's available through the API right now. But the numbers also show how far it is from the shop floor.
- Google describes the model as a high-level brain for robots: multi-step planning, spatial reasoning, coordination across several robots.
- It tracks its own progress through video and self-corrects; by their own measurements it finds the right frame for a key event with 91.3% accuracy and under a second of latency.
- Available through the API and the development environment; the enterprise platform is in closed preview.
Robotics stopped being about the hand a long time ago. It's about the head. What's new in this model is that the head watches the footage of its own work.
Hold the enthusiasm for a second. The number everyone will quote is 91.3 percent - the accuracy at spotting the key frame. The number that says more is the other one: under 60 percent accuracy at judging how far a task has gotten. On a shop floor that means on almost every other attempt the machine isn't sure it's done. Great for a demo. Not for a shop floor.
The part that genuinely pleases me is stopping when it spots a person nearby. It's not flashy and won't make a headline, but it's exactly what decides whether a robot like this can stand in a room with people. Everything else is performance. This is a permit.
This is a good step toward robots that check themselves, and it's presented honestly - the numbers are published, including the weak one. The real test, though, is on a floor that's greasy, under bad light, with nobody clearing objects out of the way.