DINO-X
DINO-X(开放世界检测)AdvancedIDEA Research’s open-world detection model that can box objects from text, examples, or no prompt at all.
DINO-X is an object-centric vision model released by IDEA Research (the Guangdong-Hong Kong-Macao Greater Bay Area Institute of Digital Economy) in November 2024, with an architecture built on Grounding DINO 1.5’s Transformer encoder-decoder. It supports text prompts, visual prompts (giving example boxes), and customized prompts, and its “universal object prompt” mode lets the model box every object in an image with no prompt at all. It was trained on Grounding-100M, a set of over 100 million grounding samples the team curated. Beyond the detection head, it also carries segmentation, keypoint, and object-captioning heads, so it can output boxes, masks, poses, and text descriptions all at once. The paper reports DINO-X Pro reaching 56.0 AP on zero-shot COCO detection; there’s also a DINO-X Edge version for edge devices. It’s mainly offered through an API, and in robotics it can be used to find a target from a single instruction, then hand off to segmentation and grasping modules.
ExampleSending a tabletop photo and the prompt “cup” to the DINO-X API returns a detection box and confidence score for every cup; feeding those boxes to SAM 2 gives pixel-level masks, which combined with depth can compute each cup’s position for grasping.
- Also called
- DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding, DINO-X Pro, DINO-X Edge
- Related
- Grounding DINO · Open-Vocabulary Object Detection · DETR · Object Detection · Segment Anything Model · Grounded SAM
- Sources
- arXiv 2411.14347: DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
IDEA-Research/DINO-X-API (GitHub) - As of
- 2025-07