Embodied AI Glossary中文

Spatial Reasoning

空间推理Common

The ability to judge an object's position, distance, size, orientation, and relation to other objects.

Spatial reasoning means a model answers questions like “is the cup to the left of the plate” or “how far apart are these two objects” based on an image or video, covering relative direction, metric distance, size comparison, and viewpoint changes. It is a prerequisite for a robot to ground a language instruction in an actual position and action. Multimodal large models are good at recognizing “what” something is but generally weaker at “where” and “how far.” Google's SpatialVLM (2024) used an automated pipeline to generate 2 billion spatial question-answer pairs with metric information from 10 million real images to train a VLM; “Thinking in Space” (2024, by Fei-Fei Li, Saining Xie, and colleagues) proposed VSI-Bench, with more than 5,000 questions, and found that current models fall clearly short of humans, while having a model first explicitly draw a “cognitive map” improves its distance judgments.

ExampleGiven the instruction “put the red cup closest to you on the right side of the bowl,” a model first has to estimate each cup's distance from the robot, then work out which region of the image corresponds to the right side of the bowl.

Also called
Spatial Understanding
Related
Spatial Intelligence · Embodied Reasoning · SpatialVLM · VSI-Bench · Vision-Language Model · 3D Visual Grounding
Sources
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces (VSI-Bench)
As of
2024-12

See it in the full glossary →