Embodiment Gap
本体差异CommonThe difference in shape, structure, and way of moving between different robots, or between humans and robots.
Embodiment refers to an agent's body: its appearance, degrees of freedom, kinematic structure, the form of its hand or gripper, sensor placement, and control interface. The embodiment gap is the set of differences between two such bodies — swap to a different robot arm and the number of joints and the action space may both change; a human hand has five fingers, while many robots have only a two-finger gripper. This gap is the central obstacle to cross-embodiment learning and to using human videos as training data: data collected, or a policy trained, on body A often cannot be used directly on body B. Common countermeasures include a unified action space or a dedicated output head per embodiment; retargeting, which maps human hand motion onto robot joints; and visual substitution, such as Mirage, which erases the target robot and replaces it with the source robot in the image, or Phantom (CoRL 2025), which erases the human arm in a human video and overlays a rendered robot arm instead.
ExampleTraining a robot on first-person video of a human hand picking up a cup: the training footage shows a human hand, but at deployment the camera sees a metal gripper — a thin cup handle a human hand can pinch may not be graspable the same way by a two-finger gripper.
- Also called
- Cross-embodiment Gap
- Related
- Cross-Embodiment · Embodiment · Unified Action Space · Motion Retargeting · Human Video Data · Cross-Painting
- Sources
- Phantom: Training Robots Without Robots Using Only Human Videos (arXiv 2503.00779)
Mirage: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting (arXiv 2402.19249)