Embodiment-agnostic
本体无关AdvancedA method or representation not tied to one particular robot's structure, usable across different robots.
“Embodiment-agnostic” describes a model, representation, or data format that is not tied to a specific robot embodiment, which can differ in number of joints, gripper type, camera placement, and control frequency. In robot learning, every robot's action space is different, so training directly on a mix of them is difficult; researchers therefore often describe “what to do” in an embodiment-agnostic form — how an object should move, as an object trajectory or scene flow, the trajectory of a hand or end effector, or a latent action — and then convert this into joint commands for a specific robot. Work by Tang and colleagues (2024) first generates a 3D scene flow for object parts, then solves for the corresponding action trajectories on different robots, and can even learn from human videos this way. Another approach has a single network directly ingest data from many embodiments at once, as CrossFormer does, controlling single arms, dual arms, wheeled vehicles, quadrotors, and quadrupeds all with one set of weights. This concept relates to cross-embodiment, unified action spaces, and embodiment-specific heads, all aimed at letting data and models be reused across different robots.
ExampleThe same object-motion trajectory, 'move the cup to the left of the plate,' can be converted separately into gripper actions for a robot arm and hand actions for a humanoid.
- Related
- Cross-Embodiment · Embodiment Gap · Unified Action Space · Embodiment-specific Head · Latent Action · Intermediate Representation
- Sources
- Embodiment-Agnostic Action Planning via Object-Part Scene Flow (Tang et al., 2024)
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation (CrossFormer)