Embodied AI Glossary中文

DINO-X

DINO-X(开放世界检测)Advanced

IDEA Research’s open-world detection model that can box objects from text, examples, or no prompt at all.

DINO-X is an object-centric vision model released by IDEA Research (the Guangdong-Hong Kong-Macao Greater Bay Area Institute of Digital Economy) in November 2024, with an architecture built on Grounding DINO 1.5’s Transformer encoder-decoder. It supports text prompts, visual prompts (giving example boxes), and customized prompts, and its “universal object prompt” mode lets the model box every object in an image with no prompt at all. It was trained on Grounding-100M, a set of over 100 million grounding samples the team curated. Beyond the detection head, it also carries segmentation, keypoint, and object-captioning heads, so it can output boxes, masks, poses, and text descriptions all at once. The paper reports DINO-X Pro reaching 56.0 AP on zero-shot COCO detection; there’s also a DINO-X Edge version for edge devices. It’s mainly offered through an API, and in robotics it can be used to find a target from a single instruction, then hand off to segmentation and grasping modules.

ExampleSending a tabletop photo and the prompt “cup” to the DINO-X API returns a detection box and confidence score for every cup; feeding those boxes to SAM 2 gives pixel-level masks, which combined with depth can compute each cup’s position for grasping.

Also called
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding, DINO-X Pro, DINO-X Edge
Related
Grounding DINO · Open-Vocabulary Object Detection · DETR · Object Detection · Segment Anything Model · Grounded SAM
Sources
arXiv 2411.14347: DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
IDEA-Research/DINO-X-API (GitHub)
As of
2025-07

See it in the full glossary →