Robots Just Started Working in Homes They've Never Seen Before
Figure's new Helix 2.5 model made beds and folded towels in 30 unfamiliar homes. What matters isn't the robot's dexterity, but how well pretraining on human behavior data transfers to new places.
For people, housework isn't something you relearn from scratch every time you enter a new home. Even in a house you've never set foot in before, you can spot the bed, spread the blanket, and pick objects up off the floor and put them away. Until now, that basic ability has been one of the hardest problems for robots to crack.
That's exactly what makes Figure's newly unveiled Helix 2.5 significant. Figure deployed the model in 30 homes it had never seen before, scattered across the San Francisco Bay Area, and had it perform three tasks: tidying a living room, folding towels, and making a bed. There was no prior data collection or environment-specific fine-tuning at the evaluation sites. Figure reported an overall task success rate of 56%.
Understanding "zero-shot" precisely
The word most likely to be overstated in this announcement is "zero-shot." It does not mean Helix 2.5 suddenly learned to fold towels or make beds with no training at all.
Figure trained the model on task-specific data for all three chores, collected at other locations. What was zero-shot was the evaluation site and its objects. The robot never saw the layouts, toys, towels, or bedding of the 30 evaluation homes during training, and it made no on-site adaptation after arriving.
That distinction matters. Many robotic systems to date have needed to collect and tune data again at each actual deployment site. That approach can work for a single factory floor, but it doesn't scale to millions of homes. If every house requires retraining, the economics of a general-purpose home robot don't hold up.
From 9% to 56% - the most important number in this announcement
Figure ran a separate comparison to isolate the effect of pretraining on its Index dataset. Model architecture, task-specific data, training method, and evaluation conditions were held constant. Only the starting state changed.
A model trained from scratch, without Index pretraining, scored a 9% zero-shot success rate. A model that had first been pretrained on Index scored 56% overall. The success bar wasn't loose, either. The living-room task required putting all 13 to 15 toys into a basket; the towel task required folding and placing the towel in a basket; the bed-making task had to meet a defined standard in full. Partial credit did not count as success.
If this result holds up, it could shift the center of gravity in the humanoid robotics race. Better hands or stronger motors alone won't be enough. The key asset may become how much, and how broadly, a company can collect diverse human behavior and fold it into a single physical world model.
Could what happened with LLMs now start happening with robots?
Figure's bigger claim is a scaling law. The company says it trained four models by scaling up Index pretraining data across an 8x range, and that prediction loss for subsequent robot behavior decreased smoothly as data increased. Figure says results from small-scale experiments alone predicted the loss of its largest training run to four decimal places, with a prediction error of just 0.54% of the total range of variation.
This is arguably the biggest industry implication of the announcement. Language models justified enormous investment on the back of an empirical scaling law: performance improves in a predictable pattern as data and compute increase. If a similar relationship holds for robots, humanoid development could look less like hands-on engineering tuning and more like a data-center-scale training race.
Figure says its Index dataset is now generating roughly 35 minutes of new human experience data every second. Converted to a daily rate, that's about 5.75 years of footage per day. To train Helix, the company is pursuing a deployment of up to 100,000 NVIDIA Vera Rubin GPUs with Nscale, and has committed an initial $3.5 billion in compute usage.
The core picture, in one frame
The old way robots scaled: New location → collect on-site data → retrain → test → deploy
What Helix 2.5 is aiming for: Broad pretraining on human behavior → define the task once → deploy directly in an unfamiliar location
If this architecture actually works, it changes the software economics of robotics substantially. It moves the industry away from a structure where training costs repeat for every new customer, and toward one where the cost of a single foundation model gets spread across a huge number of robots and locations.
But it's too early to call this the "ChatGPT moment" for home robots
A 56% success rate is meaningful at the research stage, but a consumer product faces an entirely different bar. It's hard to hand a robot that can fail more than one time in three the job of putting away glassware or handling a hot pot off the stove.
This was also an in-house experiment designed and run by Figure itself. Thirty homes is a wider range of environments than before, but three tasks aren't enough to prove general physical intelligence. Key commercialization metrics such as per-task success rates, average completion time, the frequency of human safety interventions, and whether remote assistance was used are not fully addressed by this announcement alone.
So reading this result as "the humanoid problem is solved" goes too far. The more accurate read is that the path from robots that must be retaught in every new environment to robots that transfer knowledge from large-scale pretraining on human experience has, for the first time, produced a fairly concrete set of numbers.
Insight Times Editorial Desk





