A Humanoid Walked Into 30 Homes It Had Never Seen. It Finished the Chore 56% of the Time.
Figure rented 30 homes around the Bay Area, walked a 5-foot-8 humanoid through the front door, and told it to clean up. No data was collected in any of those houses. No tuning after the robot arrived.
It finished the entire job 237 times out of 420 tries. That number, 56%, is the whole story — both halves of it.
What the announcement video shows
The clips Figure published on September 17, 2026 show its Figure 03 humanoid doing three chores in ordinary rooms: tidying toys off a living-room floor into a basket, folding towels, and making beds it has never seen before. Founder and CEO Brett Adcock appears in the video calling Helix 2.5 “the most important project we’ve ever taken on at Figure.” He and AI director Corey Lynch say the robot logged at least one success in every one of the 30 houses, which is a claim about reach, not about reliability.
Figure’s own captions flag the moments worth watching: the robot stepping back to reposition itself, shifting its stance, walking around the far side of a bed to fix a botched fold. The company calls this whole-body self-correction and treats it as the clearest visible effect of its new training approach. Adcock separately posted nearly four hours of extended footage from the homes; Humanoids Daily, which reviewed it, points out that it is not a complete record of all 420 graded trials.

The 44% that did not work
Figure graded hard, which makes the failures more honest than most demo reels. A tidying trial only counted as a success if all 13 to 15 toys ended up in the basket, with a one-minute timeout per toy. A towel trial required every towel folded and placed in the basket, three minutes each. Bed making meant both pillows and both comforter corners at the top of the bed, comforter pulled smooth. Any human stepping in for safety reasons meant the trial was scored as a failure.
By task, the company reports bed making at 94/140 (67%), towel folding at 87/140 (62%) and toy tidying at 56/140 (40%). No partial credit. So a run where the robot got 14 of 15 toys into the basket sits in the failure column — which cuts both ways when you read that 56%.
Tony Zhao of rival firm Sunday Robotics pushed back publicly, calling it failing roughly half the time and arguing that useful work means generalization plus reliability. His company reports 778 successful folds out of 785 attempts for its own system. Those numbers are not comparable: Sunday counts one garment at a time, Figure counts a whole multi-object chore. But Zhao’s underlying point stands. A robot that ruins your bed twice a shift is not something you leave alone in the house.
The real result is where the skill came from
Here is the part that matters more than the video. Figure trained two policies on identical robot data, changing exactly one thing: whether the model was first pretrained on Index, its crowdsourced dataset of ordinary people filming themselves doing everyday tasks. Trained from scratch, the robot finished 9% of trials. Pretrained on human video, 56%. Same task data, same architecture, same grading.
Figure also says the new model needed half as much robot-specific training data as an older policy that had been trained inside the room where it was tested. And it reports that doubling the human-video data improved the model’s prediction accuracy so predictably that it forecast its biggest training run’s result before running it. Index is now taking in about 35 minutes of human footage per second, from more than 90,000 weekly contributors, and the company has committed $3.5 billion of compute to training.

What you can actually do with this
Nothing yet. Figure announced no consumer price and no delivery date alongside Helix 2.5, and everything here is a company-run evaluation with no independent replication. Treat the percentages as Figure’s own scorecard.
What changed is the economics of the idea. Until now, a home robot broadly had to be taught your house. If chores learned once really do carry into a stranger’s bedroom, the expensive part of a household robot stops being per-customer setup. That is worth more than any single demo. The thing to watch next is not a prettier video — it is whether anyone outside the company can put a humanoid in an unfamiliar home and get numbers anywhere near these.
