Embodied AI Glossary中文

VSI-Bench

VSI-Bench 空间智能基准Advanced

A benchmark that uses indoor videos to test how well multimodal large models understand space.

VSI-Bench comes from the paper Thinking in Space, released in December 2024 by Saining Xie’s group at NYU together with Yale and Stanford (including Fei-Fei Li), and selected as an oral presentation at CVPR 2025. It draws 288 first-person videos from three indoor scanning datasets — ScanNet, ScanNet++, and ARKitScenes — and builds more than 5,000 question-answer pairs covering 8 task types: object counting, relative distance, relative direction, object size, absolute distance, room size, order of appearance, and route planning. The authors tested 15 multimodal large models that support video input, and all scored well below human performance (humans average about 79%), mostly struggling with spatial reasoning; language-prompting techniques such as chain-of-thought actually hurt scores, while having the model first draw a “cognitive map” improved distance judgments.

ExampleShown a video panning around a living room, the model is asked “how many meters apart are the sofa and the TV” or “standing in front of the fridge facing the sink, is the stove on your left or your right,” and is scored by numeric error or multiple-choice accuracy.

Also called
Thinking in Space: Visual-Spatial Intelligence Benchmark, Thinking in Space
Related
Spatial Intelligence · Spatial Reasoning · Multimodal Large Language Model · Benchmark · ERQA · 3D Vision
Sources
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces (arXiv)
Thinking in Space 项目主页 (Chinese)
vision-x-nyu/thinking-in-space (GitHub)
As of
2025-06

See it in the full glossary →