Embodied AI Glossary中文

Robotizing Human Videos / Human-to-Robot Video Translation

人类视频机器人化(人→机视频转换)Common

Editing footage of a human doing a task so it looks like a robot doing it, turning it into usable training data.

Human video is plentiful and cheap, but it shows a human hand and arm rather than a robot arm — a large visual embodiment gap — so training a policy on it directly performs poorly. Robotizing human video uses image editing or video generation to swap the human for a robot: first estimate the hand's 3D pose to serve as an action label, then erase the human hand and arm and inpaint the background, and finally overlay a rendered robot arm or gripper following that same trajectory. Stanford's Bohg group published Phantom in 2025, training a policy deployable directly on a real robot using only such edited human video; the same group's Masquerade pretrains a visual encoder on 675,000 frames of edited video. A late-2025 method called H2R-Grounder instead uses a fine-tuned video diffusion model to generate the robot footage directly, requiring no paired human-robot data at all.

ExamplePhantom records video of a human hand sweeping several objects together on a table, erases the hand, overlays a rendered robot arm, and trains a policy directly on this edited video that deploys straight to a real robot.

Also called
Human-to-Robot Video Translation
Related
Human Video Data · Cross-Painting · Embodiment Gap · Hand Pose Estimation · Egocentric Video · Phantom
Sources
Phantom: Training Robots Without Robots Using Only Human Videos
Masquerade: Learning from In-the-wild Human Videos using Data-Editing
H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
As of
2025-12

See it in the full glossary →