Embodied AI Glossary中文

Spatial Intelligence

空间智能Essential

The ability to understand where objects are in 3D space, and to reason and act using that understanding.

Spatial intelligence was first proposed by psychologist Howard Gardner as one of several “multiple intelligences” in his 1983 book Frames of Mind, describing the ability to picture and judge spatial relationships in one's head. In AI, the term has recently been heavily promoted by Fei-Fei Li, who co-founded World Labs; in a November 2025 essay she called spatial intelligence AI's next frontier, arguing that models need to perceive, generate, reason about, and interact with the three-dimensional world, not just understand language. In embodied AI, a robot judging “how far is the cup from the plate” or “will this fit in the cabinet” both depend on spatial intelligence; benchmarks such as VSI-Bench specifically test this ability in multimodal large models, and the results still lag noticeably behind humans.

ExampleTold to “push the chair closest to the door under the table,” a robot first has to judge which chair is closest to the door in the 3D scene, then estimate whether there's enough room underneath.

Also called
Visual-Spatial Intelligence
Related
Spatial Reasoning · World Model · 3D Vision · VSI-Bench · World Labs · Physical AI
Sources
Fei-Fei Li: From Words to Worlds: Spatial Intelligence is AI's Next Frontier
Theory of multiple intelligences - Wikipedia
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
As of
2025-11

See it in the full glossary →