Embodied AI Glossary中文

CLIPort

Advanced

A language-conditioned manipulation policy that combines CLIP's semantic understanding with Transporter Networks' pixel-level spatial precision.

CLIPort is work by Mohit Shridhar, Lucas Manuelli, and Dieter Fox at the University of Washington and NVIDIA, published at CoRL 2021. It borrows the neuroscience idea of separate “what” and “where” visual pathways. The semantic pathway uses a pretrained CLIP (an image-text contrastive learning model) to understand “what to act on” — color, shape, object category; the spatial pathway uses the fully convolutional network from Transporter Networks to process RGB-D images and decide “where to pick up, where to place,” outputting pixel-level heatmaps for grasping and placement. Fusing the two lets it follow language instructions on tabletop tasks like packing a box or folding cloth without needing object poses, segmentation masks, or symbolic state. It's an early, representative example of connecting large-scale pretrained image-text models to robot manipulation, and the same authors' later PerAct extends this idea to 3D voxels.

ExampleA single multi-task policy learned to follow language instructions across 9 real tabletop tasks using just 179 paired real image-action examples.

Also called
CLIPort: What and Where Pathways for Robotic Manipulation
Related
CLIP · Transporter Networks · Language-conditioned Policy · Tabletop Manipulation · PerAct · Pick-and-Place
Sources
CLIPort: What and Where Pathways for Robotic Manipulation (arXiv 2109.12098)
CLIPort 项目主页 (Chinese)
As of
2021-09

See it in the full glossary →