Embodied AI Glossary中文

YOLO-World

Advanced

An open-vocabulary detector that detects objects in real time just from a typed category name.

YOLO-World is an open-vocabulary object detector proposed in 2024 by Tencent AI Lab, ARC Lab, and Huazhong University of Science and Technology, published at CVPR 2024. Traditional YOLO can only detect the fixed set of categories it was trained on; YOLO-World attaches a text encoder to YOLO, uses a network called RepVL-PAN to let image features and text features interact with each other, and pretrains at large scale with a region-text contrastive loss, so a user can detect any category just by typing its name. It uses a “prompt-then-detect” strategy: the user’s vocabulary is pre-encoded and re-parameterized into the network, so the text encoder doesn’t need to run again at inference time, keeping speed close to ordinary YOLO. The paper reports 35.4 AP zero-shot on LVIS and 52 FPS on a V100. In robotics it is commonly used to find a target in real time from a language instruction, handing the result off to a grasping or navigation module.

ExampleGiven the instruction “bring me the red mug,” a program sets “red mug” as YOLO-World’s vocabulary, draws a box around the mug in real time from the wrist camera feed, and hands it to a grasp-pose detection module.

Also called
Real-Time Open-Vocabulary Object Detection
Related
Open-Vocabulary Object Detection · YOLO · Grounding DINO · OWL-ViT / OWLv2 · Object Detection · CLIP
Sources
YOLO-World: Real-Time Open-Vocabulary Object Detection (arXiv)
AILab-CVC/YOLO-World (GitHub)
As of
2025-02

See it in the full glossary →