Cross-Embodiment
跨本体EssentialTraining one model on data from many different robots so it can control more than one of them.
“Embodiment” here means a robot's specific physical body: how many arms and joints it has, whether it uses a parallel gripper or a dexterous hand, which cameras it carries. Any single robot only generates a modest amount of data on its own, so cross-embodiment learning pools data from many different robots to train one policy, letting knowledge transfer between robots and letting the result be adapted or fine-tuned to a new robot. The hard part is that different robots have different observation and action spaces — dimensions, coordinate frames, and control frequencies all vary. Open X-Embodiment (2023), a collaboration across 21 institutions, pooled data from 22 robot types; policies trained on it, RT-1-X and RT-2-X, showed positive transfer, meaning data from other robots improved performance on a given one. CrossFormer (2024) trained a single policy on 900,000 trajectories from 20 embodiments, controlling arms, wheeled robots, quadrupeds, and drones all at once.
ExampleOcto is pretrained on the multi-robot Open X-Embodiment data and can then be fine-tuned in just a few hours to a new robot with a different observation and action space.
- Also called
- Cross-embodiment Transfer, Cross-embodiment Learning
- Related
- Embodiment · Embodiment Gap · Cross-Embodiment Data · Open X-Embodiment · RT-X · CrossFormer
- Sources
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation (CrossFormer) - As of
- 2024-08