Insider Brief
- Generalist AI released GEN-1.5, a robot foundation model designed to learn physical tasks from one or a few demonstrations while also generalizing to some new situations without task-specific training.
- In tests across 10 short manipulation tasks, GEN-1.5 averaged 59% success with one-shot prompting and 83% after 10 training steps using about five minutes of data, while one-step adaptation reached 66.5% on a held-out task.
- Generalist AI said the model can combine demonstrations, imitate some human actions, transfer simulated demonstrations to real robots and improvise with unfamiliar objects and tools after more than eight months of pretraining on physical-interaction data.
Generalist AI has released GEN-1.5, a new robot foundation model designed to learn physical tasks from one or a few demonstrations while also generalizing to some new situations without task-specific training.
The company said GEN-1.5 can learn some tasks after being shown a single three- to 12-second demonstration, without updating the model’s underlying parameters. Generalist AI calls the approach “physical prompting,” with the demonstration providing sensor and movement information that the model uses to infer what the robot should do.
“Although the tasks are simple and short-horizon, this is the first model we know for which one-shot and few-shot learning of physical skills have emerged at scale, the company wrote in a blog post detailing their research. “We view these results as a significant step towards our mission of building general intelligence for the physical world.”
GEN-1.5 also supports few-shot adaptation and what Generalist AI describes as zero-shot physical generalization. That includes applying learned behaviors to unfamiliar objects or situations, improvising new movement strategies and, in some cases, using tools that were not part of the task demonstration.
GEN-1.5 Result Highlights
In tests across 10 short manipulation tasks, Generalist AI reported the following results:
- One-shot prompting: 59% average success after a single three- to 12-second demonstration, with no additional training.
- Few-shot fine-tuning: 83% average success after 10 training steps using about five minutes of data, or roughly 50 demonstrations, per task.
- One-step adaptation: 66.5% success on a held-out task after one training step using one minute of data.
The tests included opening jars, unzipping a pencil pouch, retrieving money from a purse and sweeping objects with a brush. Generalist AI noted that the tasks were relatively simple and short and that one-shot skills remain less reliable than fine-tuned models.
The model processes video, language, sensor and robot-position information and maintains about 30 seconds of context. Generalist AI said GEN-1.5 can combine separate demonstrations into longer behaviors and, in some cases, reproduce a task after watching a person perform it with their hands.
The company noted it also demonstrated zero-shot transfer from simulation to the physical world and the company reported a demonstration recorded entirely in simulation was used as a prompt for a real robot, even though GEN-1.5’s pretraining contained no simulation data and the model had not been trained on that particular task in either setting.
Generalist AI said GEN-1.5 has been pretrained for more than eight months on physical-interaction data collected from homes, warehouses, factories and other environments, with the goal of reducing the amount of task-specific data and training needed before a robot can attempt a new skill.
Generalist AI said it did not initially set out to build a model capable of learning from a single demonstration. Instead, the team focused on developing the underlying pretraining system and data pipeline needed to train robotics models on large amounts of physical interaction data.
“GEN-1.5 is a milestone we believe to be profound scientifically, not because of higher success rates, but because it represents a new frontier of generality — one that challenges our own understanding of how these models behave when pretrained at a scale of physical interaction data few thought possible without shortcuts,” the company said.