Helix 2.5 is pretrained on a vast dataset of human behaviour, and from one model it yields three skills: tidying a living room, folding towels and making a bed. Per Figure, pretraining raises success in an unfamiliar home from 9 to 56 per cent.
- Tested in 30 homes in the Bay Area, with not a single recording collected there.
- Success counts only when the task is fully completed, with no partial credit.
- Figure itself writes that general humanoid robotics is not solved.
In other people's houses no two beds are alike. A different height, a different duvet, pillows someone arranges in their own way.
A person does not relearn for every bed. Robots so far did, place by place, with data collected where they would work.
What 56 means
Fifty-six per cent sounds like a weak grade until you see how strict it is: every toy in the basket, every towel folded, the whole bed made, no partial credit. If you have ever watched a child make their bed, you know that 56 per cent is a lot.
The more important number is 9. Without pretraining the same robot with the same data succeeds in one try out of eleven, with it - in more than half. In other words most of the knowledge does not come from the task, but from watching how people live.
The hundredth home will say more than the thirtieth, and half success is not yet a robot for every living room. But for the direction they now have a curve, even if it measures prediction error, not success in the home.