Embodied AI Glossary中文

SpatialVLM

Advanced

A method that teaches a vision-language model to estimate distance and size using a massive, automatically generated set of 3D spatial question-answer pairs.

SpatialVLM was released by Google DeepMind together with MIT and Stanford in January 2024, published at CVPR 2024. The authors found that vision-language models (VLMs) can recognize what's in an image but struggle with quantitative spatial questions like 'how far is the cup from the box' or 'which one is taller,' because their training data lacks 3D spatial knowledge. They built an automated data pipeline: real photos go through object detection, segmentation, and metric depth estimation, lifting 2D images into 3D point clouds with real-world scale, and template-based spatial question-answer pairs are then generated from them — 2 billion question-answer pairs from 10 million images, the first internet-scale dataset for metric spatial reasoning. Models trained on this data show clear gains on both qualitative and quantitative spatial questions, can chain with a large language model for multi-step spatial reasoning, and can use the resulting distance estimates as a dense reward for robot tasks.

ExampleAsked 'roughly how many centimeters is the red block from the blue bowl,' an ordinary VLM typically gives only a vague description, while SpatialVLM can give a direct distance estimate with units.

Also called
Spatial VLM, SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Related
Spatial Reasoning · Vision-Language Model · Visual Question Answering · Monocular Depth Estimation · Spatial Intelligence · Google DeepMind
Sources
SpatialVLM (arXiv 2401.12168)
SpatialVLM project page
As of
2024-06

See it in the full glossary →