Embodied AI Glossary中文

Visual SLAM

视觉SLAMvSLAMCommon

Uses only cameras to simultaneously estimate where it is and build a map of the surroundings.

Visual SLAM is the branch of simultaneous localization and mapping (SLAM) that uses a camera as the main sensor — monocular, stereo, or RGB-D — often fused with an IMU (inertial measurement unit). Systems are usually split into a front end and a back end: the front end extracts feature points from images and matches them across frames to estimate camera motion (this part is also called visual odometry); the back end uses graph optimization or bundle adjustment to reduce accumulated error, and loop closure — recognizing that the camera has returned to a place it has already been — removes drift. Cameras are cheap and information-rich, but a monocular camera can’t recover absolute scale, and weak texture, reflections, and fast motion make tracking easy to lose. Well-known open-source systems include ORB-SLAM3, which supports monocular, stereo, RGB-D, and visual-inertial modes, and the learning-based DROID-SLAM; visual SLAM is widely used for localization and navigation in mobile robots, AR glasses, and drones.

ExampleThe ORB-SLAM3 paper reports that, running on the EuRoC drone dataset with stereo cameras plus an IMU, the average error of the estimated trajectory is about 3.6 centimeters; the input is an image sequence, and the output is the camera pose for every frame plus a sparse map point cloud.

Also called
vSLAM, VSLAM
Related
Simultaneous Localization and Mapping · Visual Odometry · Visual-Inertial Odometry · Loop Closure Detection · ORB-SLAM3 · Absolute Trajectory Error / Relative Pose Error
Sources
MathWorks: What Is SLAM (Simultaneous Localization and Mapping)?
ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM (arXiv 2007.11898)

See it in the full glossary →