Embodied AI Glossary中文

Pointing

指向(点预测)Common

A vision-language model answering by marking exact pixel coordinates on the image, following a text instruction.

Pointing lets a vision-language model (VLM) answer by literally marking a spot on the image: given an image and an instruction, it outputs one or more 2D pixel coordinates. Ai2's Molmo (2024) specifically collected the PixMo-Points dataset for this and can point at and count objects; RoboPoint (2024) fine-tunes a VLM on automatically synthesized data to predict keypoints for where to grasp or where to place something; Google DeepMind's Gemini Robotics-ER outputs coordinates in [y, x] format, normalized to a 0–1000 range. A point is more fine-grained than a box and more directly usable downstream than plain text: combined with back-projection from a depth map, it gives a 3D target position that can be handed directly to grasping, motion planning, or a VLA model.

ExampleAsking Gemini Robotics-ER to ‘point to all the bananas in the image’ returns a list of entries like point: [376, 508], label: small banana, with each point marking one banana's location.

Also called
2D Point Prediction, Point Prediction
Related
Molmo (Ai2) · RoboPoint · Gemini Robotics-ER · Visual Prompting · Affordance Detection · Projection / Back-Projection
Sources
Molmo and PixMo (arXiv 2409.17146)
RoboPoint (arXiv 2406.10721)
Gemini API Docs: Gemini Robotics-ER overview
As of
2026-09

See it in the full glossary →