Embodied AI Glossary中文

Grounding DINO

Common

An open-vocabulary detector that finds and boxes whatever object a text description names, given an image and text.

Grounding DINO is an open-set object detection model released in March 2023 by IDEA Research together with Tsinghua University and other collaborators, with the paper later accepted at ECCV 2024. Traditional detectors only recognize the few dozen fixed classes they were trained on; Grounding DINO instead combines the transformer-based detector DINO with a text encoder, performing multi-layer fusion between image and text features, so a user can input a class name or a short description (such as ‘red cup’) and get back matching detection boxes along with the matched words. The paper reports 52.5 AP in zero-shot detection on COCO without using any COCO training data. The code is open-sourced under Apache 2.0 and has been integrated into Hugging Face Transformers. It's often chained with the segmentation model SAM into a pipeline called Grounded-SAM — box by text first, then cut a mask — a common front end for robots that need to ‘find things by instruction.’ Later versions, Grounding DINO 1.5 and 1.6, are only available through an API.

ExampleA user says ‘put the banana in the bowl.’ The system runs Grounding DINO with the prompt ‘banana. bowl.’ to box both objects, uses SAM to get their masks, and combines that with the depth map to compute the grasp point and the placement point.

Also called
GroundingDINO
Related
Open-Vocabulary Object Detection · Grounded SAM · Segment Anything Model · OWL-ViT / OWLv2 · YOLO-World · DINO-X
Sources
Grounding DINO (arXiv:2303.05499)
IDEA-Research/GroundingDINO (GitHub)
IDEA-Research/Grounding-DINO-1.5-API (GitHub)
As of
2024-07

See it in the full glossary →