Embodied AI Glossary中文

Phantom

Phantom(无机器人训练)Advanced

A method that trains robot policies purely from human demonstration videos by digitally replacing the human hand with a rendered robot arm.

Phantom comes from Jeannette Bohg's lab at Stanford University (Marion Lepert and colleagues), released in March 2025 and presented at CoRL 2025. Teleoperated data collection is expensive, while human videos are cheap — but a human hand looks nothing like a robot gripper, and the videos carry no robot action labels. Phantom's solution: estimate hand pose in every frame and convert it into robot end-effector actions, then use image inpainting to erase the human hand from each frame and render a virtual robotic arm in its place, so the training images look like a robot performing the task. This lets a policy be trained with zero robot data and deployed zero-shot on a Franka or a Kinova Gen3 arm, completing tasks such as pick-and-place, cup stacking, tying rope, sweeping, and insertion, with a reported success rate as high as 92%. It belongs to the 'human-video-to-robot' family of methods.

ExampleA researcher demonstrates 'stack the cups' with their own hand on a tabletop; Phantom automatically replaces the hand in the video with a rendered gripper and generates action labels, and the resulting policy runs directly on a Franka arm.

Also called
Phantom: Training Robots Without Robots Using Only Human Videos
Related
Robotizing Human Videos / Human-to-Robot Video Translation · Human Video Data · Robot-free (Embodiment-free) Data Collection · Hand Pose Estimation · EgoMimic · Embodiment Gap
Sources
arXiv 2503.00779: Phantom
Phantom 项目主页 (Chinese)
As of
2025-09

See it in the full glossary →