Embodied AI Glossary中文

Semantic SLAM

语义SLAMAdvanced

SLAM that recognizes object categories while localizing and mapping, producing a map with semantic labels attached.

Semantic SLAM is a family of methods that adds semantic information to traditional SLAM (Simultaneous Localization and Mapping): while estimating its own pose and building a geometric map, the system also uses a detection or segmentation network to recognize walls, tables, chairs, and other objects, and writes those categories into the map. Semantics can also help SLAM itself — for instance, by filtering out moving objects such as people and cars to reduce their interference with pose estimation, or by using objects as landmarks for loop closure detection. Representative work includes SLAM++ (2013), which builds maps at the level of individual objects, and Kimera, open-sourced by MIT’s SPARK Lab in 2019, which uses a camera and IMU to generate a semantically labeled 3D mesh in real time on a CPU. Recent work has also explored semantic SLAM based on NeRF and 3D Gaussian Splatting, though the computational cost is still high for embedded platforms. Its output is, in effect, a semantic map.

ExampleKimera walks a stereo camera and IMU around an indoor space and outputs, in real time, a 3D mesh map labeled with semantics such as “wall,” “floor,” and “chair.”

Also called
Metric-Semantic SLAM
Related
Simultaneous Localization and Mapping · Semantic Map · Visual SLAM · 3D Scene Graph · Gaussian Splatting SLAM · ConceptGraphs
Sources
arXiv 1910.02490: Kimera
arXiv 2505.12384: Is Semantic SLAM Ready for Embedded Systems?

See it in the full glossary →