Open-Vocabulary Object Detection
开放词汇检测OVDCommonDetecting whatever object a piece of text names, instead of being limited to a fixed list of trained classes.
Traditional detectors only recognize the fixed set of classes they saw box-labeled at training time (COCO's 80, for instance). Open-vocabulary detection instead lets a detector accept an arbitrary text description and box classes it never saw with box labels during training. The setting was formally proposed by Alireza Zareian and colleagues at CVPR 2021: box annotations from a small set of base classes are combined with vision-language alignment learned from large-scale image-text pairs to detect new classes. OWL-ViT, Grounding DINO, and YOLO-World all followed this path afterward, with Grounding DINO reaching 52.5 AP in zero-shot detection without any COCO training. This is very practical for robots: if a user says ‘bring me the blue mug,’ the system can box it directly from that sentence. Strictly speaking, ‘open-set detection’ originally meant recognizing unseen objects as simply ‘unknown,’ but the term is now often used interchangeably with open-vocabulary detection.
ExampleIn the Grounded-SAM pipeline, Grounding DINO first boxes the target using the text ‘banana,’ then the box is handed to SAM to get a pixel-level mask, and the robot computes a grasp point from that.
- Also called
- OVD, Open-Set Detection, Text-Guided Detection
- Related
- Open-vocabulary · Grounding DINO · OWL-ViT / OWLv2 · YOLO-World · Grounded SAM · Object Detection
- Sources
- Open-Vocabulary Object Detection Using Captions (CVPR 2021)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection