Embodied AI Glossary中文

Semantic Map

语义地图Advanced

A map that labels what each place is, on top of its geometry, so it can be queried by object name or natural language.

An ordinary SLAM map only records where obstacles are and where surfaces sit; a semantic map goes further, attaching category labels or features to the points, cells, or objects in the map, so it can answer questions like “where is the refrigerator” or “where is the kitchen.” Early approaches labeled the map using detection or segmentation networks with a fixed set of categories; in recent years, open-vocabulary semantic maps have become popular, storing features from vision-language models such as CLIP or LSeg directly in the map so it can be queried with arbitrary natural language. A representative example is VLMaps (2022, from the University of Freiburg, Google, and others), which extracts pixel-level vision-language features from RGB-D video, back-projects them into 3D using depth, and stores them in a top-down grid map, letting it follow navigation instructions with spatial relationships such as “go between the sofa and the TV.” Semantic maps can also be organized as 3D scene graphs, and are a common intermediate representation for object-goal navigation, vision-and-language navigation, and mobile manipulation.

ExampleVLMaps lets a LoCoBot and a drone share the same language map, navigating from instructions such as “move three meters to the right of the chair.”

Also called
Open-Vocabulary Semantic Map, Language Map
Related
Semantic SLAM · VLMaps · ConceptGraphs · 3D Scene Graph · Object-Goal Navigation · Occupancy Grid Map
Sources
VLMaps 项目主页 (Chinese)
arXiv 2210.05714: Visual Language Maps for Robot Navigation

See it in the full glossary →