Insider Brief
- Figure introduced Helix 2.5, a neural network that enabled its humanoid robot to perform three whole-body household tasks across 30 unseen homes without collecting data or adapting the model in those environments.
- Figure said Index pretraining raised zero-shot success from 9% to 56%, while Helix 2.5 used half as much task-specific data as a comparable Helix 02 behavior and expanded across 30 homes.
- The company reported predictable gains in robot-action prediction as Index pretraining data increased, while Index now generates about 35 minutes of human-experience data per second and Figure has committed $3.5 billion in compute to Helix.
Figure has introduced Helix 2.5, a new neural network designed to let its humanoid robots perform whole-body household tasks in unfamiliar homes without collecting training data in those locations.
“The holy grail for robotics is being able to generalize: doing work in unseen places,” founder and CEO Brett Adcock wrote in a LinkedIn post. “We rented 30 homes in the Bay Area and are doing tasks without any new training. Helix 2.5 was built to answer a harder question: can a humanoid enter a home it has never seen and immediately get to work, with its whole body, on its own?”
Figure pointed out in a blog post announcing the new foundation model that current robots typically learn the environments where they operate one location at a time and Helix 2.5 was built to test whether a humanoid could instead transfer knowledge learned from human behavior into new rooms, layouts and objects without fine-tuning or adaptation in those environments.
The Tests
The company reported testing Helix 2.5 across 30 Bay Area homes that were not included in its training data. Using a foundation model pretrained on Figure’s Index dataset of human behavior, the humanoid performed the three tasks of tidying living rooms, folding towels and making beds. No data was collected in the evaluation homes, and the objects used in testing did not appear in the task-specific training data.
- Living-room tidying: Pick up all 13 to 15 toys scattered around the room and place them in a basket.
- Towel folding: Pick up, fold and place all towels in a basket.
- Bed making: Position both pillows near the top of the bed, pull the comforter corners into place and smooth the comforter.
Success required the robot to complete the entire task. For living-room tidying, it had to place all 13 to 15 toys in a basket. For towels, every towel had to be folded and placed in the basket. Bed making required both pillows and the comforter corners to be positioned at the top of the bed, with the comforter pulled smooth. The same model checkpoint was used across all 30 homes.
The Results
Figure said Helix 2.5 cut the amount of task-specific data needed to specify a behavior in half while expanding its scope from a single environment to 30 unseen homes. The company described that as making behavior specification 2x cheaper while increasing its deployment scope 30x.
Figure said Index pretraining accounted for most of Helix 2.5’s zero-shot capability. In a controlled comparison using identical task data and model architecture, the Index-pretrained policy succeeded on 56% of trials, compared with 9% for the model trained from scratch. No single evaluation task represented more than 1.9% of the Index pretraining dataset, the company said.
Figure also reported improved self-correction during longer tasks. The robot could step back, reposition its body, change its stance or move around a bed to recover from mistakes and continue the task.
Human-to-Robot Transfer
The company said it separately tested how performance changed as more human-behavior data was added. According to Figure, across four models trained on progressively larger subsets of Index, robot-action prediction improved predictably with each doubling of pretraining data. Figure noted the smaller training runs were accurate enough to forecast the largest run’s test loss to four decimal places before training began.
“Loss fell predictably with each doubling of Index,” the company noted in its blog post. “To our knowledge, this is the first human-to-robot transfer scaling law measured on a humanoid: Index pretraining predictably improves downstream next robot action prediction.”
What’s Next?
Figure pointed out that Index is generating about 35 minutes of new human-experience data every second. The company has also committed $3.5 billion in computing resources to training Helix.
“If Helix 2.5 is any guide, more data and compute should translate directly into more of the physical world learned before the robot ever enters a new home,” the company said. “It’s time to scale up.”
Image credit: Figure