Embodied AI Glossary中文

SuSIE

Advanced

A method that uses an image-editing diffusion model as a high-level planner, sketching a subgoal image for a low-level policy to reach.

SuSIE was released by Kevin Black, Sergey Levine, Chelsea Finn, and colleagues at UC Berkeley and Stanford in October 2023. The method has two levels. The high level is InstructPix2Pix, an open-source image-editing model, fine-tuned on human video and robot data: given the current camera image and a language instruction, it outputs a subgoal image showing 'what it should look like' at some point soon. The low level is a goal-conditioned policy that looks only at the target image, not the language, and is responsible for getting the robot from the current image to that subgoal image. The two levels alternate until the task is complete, letting the high level draw on internet-scale image pretraining to handle new objects and instructions, while the low level focuses purely on precise control. It achieved state-of-the-art results at the time in the zero-shot setting on the CALVIN benchmark, and outperformed RT-2-X in real-robot experiments too. SuSIE belongs to the same 'generate an image, then derive the action' family as UniPi.

ExampleGiven the instruction 'put the yellow block in the drawer,' SuSIE first generates an image of the gripper already holding the yellow block as a subgoal; once the low-level policy reaches that state, it generates the next subgoal image showing the block inside the drawer.

Also called
Subgoal Synthesis via Image Editing, Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
Related
Goal-conditioned Policy · Diffusion Model · Hierarchical Architecture · CALVIN Benchmark · UniPi · Zero-shot
Sources
Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models (arXiv 2310.10639)
SuSIE project page
As of
2023-10

See it in the full glossary →