Embodied AI Glossary中文

COCO / LVIS

COCO / LVIS 数据集Advanced

The most widely used object-detection and segmentation benchmark; LVIS extends COCO's images to over a thousand long-tail categories.

COCO was released by Microsoft and collaborators in 2014: about 330,000 everyday-scene images with 1.5 million object instances, using 80 categories for detection and instance segmentation, plus annotations for image captioning, human keypoints, and more. It's the most common training and evaluation set for detection and segmentation models, and mAP (mean average precision) is usually reported on it. LVIS was introduced by Agrim Gupta, Piotr Dollár, and Ross Girshick at Facebook AI Research at CVPR 2019; it reuses COCO's images but re-annotates them with 1,203 categories and about 2 million instance masks, following a long-tail distribution where a few categories are common and most have very few examples, specifically to test recognition of rare categories. Open-vocabulary detection (finding objects from an arbitrary text description) commonly uses LVIS's rare classes to measure zero-shot ability. Most detection and segmentation models used in robot perception are pretrained or evaluated on one or both of these datasets.

ExampleNew releases in the YOLO family typically report mAP on the COCO val2017 split; open-vocabulary detectors like YOLO-World and Grounding DINO instead report zero-shot AP on LVIS.

Also called
Common Objects in Context, Large Vocabulary Instance Segmentation, MS COCO, LVIS v1.0
Related
Object Detection · Instance Segmentation · Open-Vocabulary Object Detection · Mean Average Precision · YOLO · Grounding DINO
Sources
Microsoft COCO: Common Objects in Context (arXiv 1405.0312)
LVIS: A Dataset for Large Vocabulary Instance Segmentation (arXiv 1908.03195)
COCO 官网数据集介绍 (Chinese)

See it in the full glossary →