Embodied AI Glossary中文

Open-vocabulary

开放词汇Common

A model that can recognize or handle objects and concepts described in arbitrary words, beyond its training categories.

Traditional detection and segmentation models only recognize a fixed set of categories seen during training, such as COCO's 80 classes — this is called closed-vocabulary. Open-vocabulary means a model can accept any text as a category, including words that never appeared in its training annotations. Alireza Zareian and colleagues proposed open-vocabulary object detection with OVR-CNN in 2020: first learn a shared visual-semantic space from large amounts of image-text pairs, then train a detector with only a small number of bounding-box annotations. Contrastive image-text models like CLIP later popularized this approach, leading to systems such as OWL-ViT, Grounding DINO, and YOLO-World. In embodied AI, open-vocabulary capability lets a robot understand object names that were never defined in advance, and it is commonly used for open-vocabulary grasping, navigation, and mobile manipulation.

ExampleOK-Robot, in a real home, accepts instructions naming arbitrary objects, such as 'put the stuffed rabbit in the basket,' first using an open-vocabulary model to locate the object before grasping and placing it.

Related
Open-Vocabulary Object Detection · Open-Vocabulary Segmentation · Zero-shot · CLIP · OK-Robot · Open-world
Sources
Open-Vocabulary Object Detection Using Captions
OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

See it in the full glossary →