Referring Expression Segmentation
指代表达分割RESAdvancedGiven a sentence describing one specific object, precisely segmenting out exactly that object in the image.
Referring expression segmentation takes an image plus a natural-language description (such as “the two people sitting on the bench on the right”) and outputs a pixel-level mask of the object the sentence refers to. It is more fine-grained than open-vocabulary segmentation: the latter segments every object of a named category, while referring segmentation must pick out one specific instance based on color, position, relationships, and other descriptive cues. Hu et al. proposed an early end-to-end method in 2016, encoding the sentence with an LSTM and fusing it with a convolutional network’s feature map to predict a mask pixel by pixel. Common benchmarks include RefCOCO, RefCOCO+, and G-Ref. Today the task is often handled by multimodal large models or grounding models paired with a SAM-style segmentation model. When a robot hears “hand me that red cup on the left,” it must first use this to find the pixel region of the target, then combine that with depth to get a 3D position for grasping.
ExampleGiven the instruction “pick up the blue block on the far left of the table,” the model outputs a mask for exactly that block in the wrist camera’s image.
- Also called
- RES, Referring Image Segmentation
- Related
- Visual Grounding · Open-Vocabulary Segmentation · Instance Segmentation · Mask · Grounded SAM · Language Grounding
- Sources
- Segmentation from Natural Language Expressions (arXiv 1603.06180)