Visual Prompting
视觉提示AdvancedDrawing boxes or numbers directly on an image so a multimodal large model can answer by referring to those marks.
Visual prompting means overlaying boxes, points, arrows, or numbers directly on an input image, without changing any model weights, to guide a multimodal large model to attend to and refer to specific regions. The representative method is Set-of-Mark (SoM), proposed by a Microsoft team in 2023: a segmentation model such as SAM or SEEM first cuts the image into regions, each region is labeled with a number, mask, or box, and GPT-4V is then asked to answer using those numbers; the paper reports that this beats fully fine-tuned specialist models zero-shot on the RefCOCOg referring task. It addresses the fact that large models can describe a scene in words but struggle to state a precise pixel location in text. In robotics, work such as MOKA and PIVOT uses this kind of marking to let a vision-language model pick out a grasp point or a direction to move, which is then handed to low-level control to execute.
ExampleObjects in a tabletop photo are labeled 1 through 8; asked “which object could be used to scoop soup,” the model answers “number 5,” and the program converts region 5 into 3D coordinates for the robot arm.
- Also called
- Set-of-Mark, SoM, Marker-Based Visual Prompting
- Related
- Visual Prompting (Set-of-Mark) · MOKA · PIVOT · Visual Grounding · Segment Anything Model · Vision-Language Model
- Sources
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V (arXiv)
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting (arXiv)