Tech

A Robot Just Cleaned Houses It Had Never Seen Before

Figure's new Helix 2.5 model made beds and folded towels in 30 unfamiliar homes. What matters more than the robot's hands is how well pretraining on human behavior data transfers to new places.

For humans, housework does not mean starting from scratch every time you enter a new place. Walk into an unfamiliar house and you can still spot the bed, straighten the sheets, pick items off the floor and put them away. Until now, that basic ability has been one of the hardest problems for robots to solve.

That is exactly what makes Figure's newly unveiled Helix 2.5 significant. The company sent the model into 30 homes it had never visited in the San Francisco Bay Area and had it tidy living rooms, fold towels and make beds. There was no prior data collection or environment-specific fine-tuning at any of the evaluation homes. Figure reported an overall task success rate of 56%.

Understanding "zero-shot" precisely

The word most likely to get overstated in this announcement is "zero-shot." It does not mean Helix 2.5 suddenly learned to fold towels or make beds with no training at all.

Figure trained the model on task-specific data for all three chores, collected at other locations. What was zero-shot was the evaluation site and its objects. The robot never saw the layouts, toys, towels or bedding of the 30 evaluation homes during training, and it received no additional adaptation once it arrived on site.

That distinction matters. Many robotic systems today still need to collect and tune data fresh at each deployment location. That approach can work in a single factory, but it does not scale to millions of homes. If every household requires its own retraining, the economics of a general-purpose home robot fall apart.

9% to 56%: the number that matters most here

<div class="metric-grid"> <div class="metric"><div class="num">9%</div><div class="label">Zero-shot success rate for a policy trained from scratch, without Index pretraining</div></div> <div class="metric"><div class="num">56%</div><div class="label">Overall task success rate for a policy pretrained on Index</div></div> <div class="metric"><div class="num">30</div><div class="label">Real homes evaluated with no prior data collection</div></div> </div>

Figure isolated the effect of Index pretraining specifically. Model architecture, task-specific data, training method and evaluation conditions were held constant, and only the starting state was changed.

The model trained from scratch scored 9%. The model pretrained on Index scored 56%. The success criteria were not lenient, either. Tidying the living room required putting all 13 to 15 toys into a basket. Towels had to be folded and placed in a basket. Making the bed had to meet a fixed standard all the way through. Partial credit did not count as success.

If this result holds up, it could shift where the humanoid robotics race is actually won. A better hand or a stronger motor alone will not be enough. How much diverse human behavior a company can collect, and how broadly, to train a single physical world model may become the core asset.

Is what happened with LLMs now starting with robots

Figure's bigger claim is a scaling law. The company says it trained four models with Index pretraining data scaled up across an eightfold range, and that the loss for predicting subsequent robot behavior fell smoothly as the data grew. Using only the results from smaller experiments, the company says it predicted the loss of its largest training run to four decimal places, with a prediction error equal to 0.54% of the total range of change.

This is arguably the biggest industrial implication of the announcement. Language models justified enormous investment on the back of an empirical scaling law: performance improves in a predictable pattern as data and compute grow. If a similar relationship holds for robots, humanoid development could look less like manual engineering tuning and more like a data-center-scale training race.

Figure says its Index system is currently generating roughly 35 minutes of new human experience data every second, which works out to about 5.75 years of footage per day. To train Helix, the company is pursuing a deployment of up to 100,000 NVIDIA Vera Rubin GPUs with Nscale and has committed an initial $3.5 billion in compute.

The one chart that matters

How robots used to scale: New location → collect data on-site → retrain → test → deploy

What Helix 2.5 is aiming for: Broad pretraining on human behavior → define a task once → deploy directly to an unfamiliar location

If this architecture actually works, it would change the software economics of robotics substantially. Instead of a structure where training cost repeats with every new customer, the cost of a single foundation model gets spread across a huge number of robots and locations.

Still too early to call it the "ChatGPT moment" for home robots

A 56% success rate is meaningful at the research stage, but a consumer product faces an entirely different bar. A robot that can fail more than one time in three is not one you hand a wine glass or a hot pan to.

This was also a company-designed, company-run experiment. Thirty homes is a wider range of environments than earlier tests, but three tasks are not enough to prove general physical intelligence. Key commercialization metrics, such as per-task success rates, average completion time, how often a human had to step in for safety, and whether remote assistance was used, are not fully confirmed by this announcement alone.

Reading these results as proof that the humanoid robot problem is "solved" would be overreaching. The more accurate read is that the path from robots that must be retaught in every new environment to robots that transfer knowledge learned from pretraining on human experience at scale has, for the first time, produced a fairly concrete set of numbers.

Figure's Helix 2.5 jumped from a 9% to a 56% zero-shot success rate on home chores once pretrained on broad human-behavior data, offering the first concrete evidence that the humanoid robotics race may hinge on data scale rather than hardware dexterity.

Insight Times Editorial Desk