Embodied AI Glossary中文

PIVOT

PIVOT(迭代视觉提示)Advanced

A method that controls robots zero-shot by drawing candidate actions on an image and having a VLM repeatedly pick and narrow them down.

PIVOT was led by Google DeepMind (Soroush Nasiriany, Fei Xia, and 21 other co-authors), released in February 2024 and published at ICML 2024. Vision-language models (VLMs) can only output text, but robots need continuous coordinates and actions. PIVOT reframes the problem as iterative visual question answering: it samples a batch of candidates — target points, movement directions, or trajectories — and draws them on the image as numbered arrows or dots; the VLM picks the best few; a new distribution is fit around the selected candidates and resampled, narrowing the range each round, until a final action emerges after a few iterations. The whole process needs no robot training data, and works for real-robot navigation, tabletop manipulation, following instructions in simulation, and image grounding — though the authors acknowledge the success rate is still far from practical. PIVOT represents the visual-prompting line of work, alongside similar approaches like MOKA and Set-of-Mark prompting.

ExampleTo send a mobile robot to grab a soda can on a table: in round one, several numbered arrows for possible directions are drawn on the image and the VLM picks arrows 3 and 5; the next round redraws arrows only near those two directions and picks again, until the direction is precise enough.

Also called
Iterative Visual Prompting, PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Related
Visual Prompting (Set-of-Mark) · MOKA · Vision-Language Model · Zero-shot · Cross-Entropy Method · RoboPoint
Sources
arXiv 2402.07872: PIVOT
ICML 2024 论文页(PMLR v235) (Chinese)
PIVOT 项目主页 (Chinese)
As of
2024-07

See it in the full glossary →