Embodied AI Glossary中文

Cross-Painting

跨本体图像替换Advanced

Digitally erasing the deployed robot from the camera feed and painting in the robot the policy was trained on, so a vision policy transfers without retraining.

Cross-painting was introduced by UC Berkeley's Ken Goldberg lab together with Google DeepMind researchers in the Mirage method (RSS 2024). A vision-based policy that has only ever seen a “source” robot tends to fail once deployed on a differently shaped “target” robot, because the camera image looks different. Cross-painting fixes this at run time: it uses a segmentation mask to erase the target robot from the video feed and inpaint the background, then uses the robot's URDF (the file format that describes a robot's links and joints) to compute the joint angles that would place the source robot's end effector at the same pose, and renders the source robot into the scene — so the policy always “sees” the robot it was trained on. Differences in how the two robots actually move are compensated separately with a forward-dynamics model (which predicts where the end effector ends up after a given action). Mirage achieved zero-shot transfer across a Franka arm, a UR5 arm, and different grippers. A similar segment-and-render pipeline has also been used to turn human hands in video into robot hands.

ExampleA policy trained only on data collected with a Franka arm is deployed on a UR5; at run time, the UR5 in the camera feed is replaced with a rendered Franka, and the policy completes tasks like grasping with no retraining.

Also called
Mirage, Robot Embodiment Visual Swap
Related
Cross-Embodiment · Embodiment Gap · Zero-shot · Robotizing Human Videos / Human-to-Robot Video Translation · Generative Data Augmentation · Forward Dynamics Model
Sources
Mirage: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting (arXiv 2402.19249)
BerkeleyAutomation/mirage (GitHub)
As of
2024-07

See it in the full glossary →