OWL-ViT / OWLv2
AdvancedA Google open-vocabulary object detection model that can find objects from a text description or an example image.
OWL-ViT was proposed by Matthias Minderer and colleagues at Google, published at ECCV 2022. The approach first pretrains a vision transformer with CLIP-style image-text contrastive learning, then fine-tunes it end to end into a detector: each image patch outputs a box and a feature vector, which is compared against text features by similarity, allowing detection of categories never seen during training; it also supports one-shot detection from a single example image. OWLv2, from 2023, scales up the training data through self-training: an existing detector automatically generates pseudo-box labels on web image-text pairs, producing over 1 billion examples, which raised average precision on LVIS rare categories from 31.2% to 44.6%. Both models are available in Hugging Face Transformers and are commonly used in robotic systems to locate objects from language instructions.
ExampleGiven a desktop image and the text “a red mug”, OWLv2 returns a detection box and confidence score for the mug, and the robot estimates a grasp position from the depth inside that box.
- Also called
- Open-World Localization Vision Transformer, OWL-ST
- Related
- Open-Vocabulary Object Detection · CLIP · Grounding DINO · YOLO-World · Vision Transformer · OK-Robot
- Sources
- Simple Open-Vocabulary Object Detection with Vision Transformers (arXiv 2205.06230)
Scaling Open-Vocabulary Object Detection (arXiv 2306.09683)
Hugging Face Transformers: OWLv2