Article 4 min read 810 words

Figure Helix 2.5: What Unseen-Home Tests Show Buyers

A useful home-robot test should ask whether a skill survives a change of room, objects and furniture. This explainer examines that question through Figure's latest published home evaluation, then separates the research result from the evidence a household would need before buying.

ui44 Team All articles

This is source-based analysis, not a hands-on review. Our earlier Figure 03 explainer covered hardware and Helix 02 demos; the question here is transfer to unfamiliar homes.

What Figure tested

Figure's September 17, 2026 Helix 2.5 report describes toy tidying, towel folding and bed making across 30 unseen Bay Area homes. These are manufacturer-reported results, not independent testing.

“Zero-shot” concerns homes and objects; tasks were fine-tuned elsewhere. Each task used a fixed checkpoint, not selected using evaluation performance.

The report's chart shows 237/420 complete-task successes with Index human-behavior pretraining (56%), versus 35/420 without it (about 8%). The prose instead says 9% for the baseline; the two disagree. The comparison holds architecture, optimization, hyperparameters, task data and evaluation fixed; the baseline starts from random weights. It is not a competitor comparison.

Each policy had 140 trials per task. The Index results were 87 towel, 56 toy and 94 bed successes; baseline results were 12, 7 and 16 respectively. Whole-task completion was required; the appendix grades neatness separately, and the chart specifies a minimum bedding grade. The appendix imposes timeouts and counts safety interventions as failed trials.

How to interpret the result

The useful inference is about the training recipe: under this controlled comparison, prior learning helped a robot apply specified behaviors in unfamiliar settings. That is a more informative question than whether a selected video looks impressive. It still leaves a gap between performing a chore under an evaluation procedure and taking responsibility for it in your home.

Do not treat the headline percentage as your household's probability of success. Your furniture, tolerated mess, interruptions and definition of a finished job may differ. Trial counts help interpret a rate, but do not by themselves explain how results vary between homes. To judge consistency, ask for per-home breakdowns, variation across repeats and failure causes.

Read the task-level results before the pooled rate. If your priority is laundry, performance on other chores cannot answer whether your towels end up where you want them. Ask to see the exact end condition, time allowed and what happens to an interrupted attempt before comparing headline scores.

Questions for a buying decision

Use these questions when reading any home-robot evaluation. They are purchasing questions, not additional capabilities established by this study.

Your decision

Will it handle my chore?

Evidence to request
A matching task, objects and finish standard, including unsuccessful attempts.

Your decision

Will it save me time?

Evidence to request
Task duration plus setup, resets, cleanup and human assistance.

Your decision

Can I leave it working?

Evidence to request
Documented supervision requirements, stopping behavior and tests with interruptions.

Your decision

What happens after a mistake?

Evidence to request
Failure categories, recovery outcomes and conditions requiring an operator.

Your decision

Can I actually buy and support it?

Evidence to request
Written regional ordering, delivery, repair service and recurring-cost terms.

Keep these answers separate. A successful movement demonstration cannot supply a missing support commitment; an order page cannot supply missing reliability evidence. If a vendor offers an in-home trial, agree on the chore and acceptable finish beforehand, and record your own assistance time alongside the robot's running time. A task that needs you to prepare every object may still be useful, but it is a different service from cleaning up after you without preparation.

The Figure 03 profile and Figure AI manufacturer page provide starting points for checking product records and their sources. A catalog status such as Active is not a promise of consumer availability. For a comparison, retain each price's currency and check whether the listed configuration includes the demonstrated capability; do not infer software access from a hardware listing.

Questions buyers may still have

Does this establish general home autonomy?

A bounded evaluation cannot establish unrestricted household operation. Ask which chores, rooms and conditions are supported, what requires supervision, and what the system must refuse. Cooking, care work or handling valuables should not be assumed from success at another task.

Can I rank other robots against this percentage?

Only a matched protocol makes that comparison meaningful. Different objects, time limits, grading rules and intervention policies can change the question being measured. Put unmatched results side by side as evidence descriptions, not a performance leaderboard.

Is this a retail announcement or independent review?

The cited evidence is a manufacturer research report. A buying decision needs separate, current commercial terms and evidence from use outside the supplier's evaluation. ui44 did not operate the robot or reproduce the experiment.

Source and update note

Source

Figure's Helix 2.5 report, results chart and appendix

Evidence type
Primary; manufacturer research
Dates
Published September 17, 2026; checked October 10, 2026

Revisit this analysis when Figure reconciles the baseline percentage, or when independent replication or consumer delivery terms provide new evidence.

UT

Written by

ui44 Team

Published October 10, 2026

Share this article

Open a plain share link on X or Bluesky. No embeds, no widgets, no cookie baggage.

Explore the database

Go beyond the headlines

Compare specs, features, and prices across 100+ robots from leading manufacturers worldwide.