Embodied AI Glossary中文

DROID-SLAM

Advanced

A deep-learning visual SLAM system from Princeton that iteratively estimates camera pose and dense depth.

DROID-SLAM is a visual SLAM (simultaneous localization and mapping) system from Princeton University’s Zachary Teed and Jia Deng, published at NeurIPS 2021. It borrows the architecture of the same group’s RAFT optical flow model, computing dense pixel correspondences between correlated frames, and uses a recurrent network to repeatedly update the camera poses and the depth of every pixel, with a differentiable dense bundle adjustment (BA) layer embedded in the middle — jointly optimizing camera poses and 3D structure — putting geometric constraints directly inside the network. It is trained purely on monocular video from the synthetic TartanAir dataset, but at test time it can also take stereo or RGB-D input, and it clearly outperforms earlier methods in accuracy on TartanAir, EuRoC, TUM-RGBD, and ETH3D, with far fewer catastrophic failures. The cost is dependence on a GPU — inference needs at least 11 GB of VRAM. Later deep-SLAM work like MASt3R-SLAM often uses it as a comparison baseline.

ExampleGiven a video of someone walking around a room with a handheld camera, DROID-SLAM outputs the camera pose and dense depth for every frame, which can be stitched into a point cloud of the room, or used to recover the camera trajectory from a human demonstration video.

Also called
DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
Related
Visual SLAM · Bundle Adjustment · RAFT · ORB-SLAM3 · MASt3R-SLAM · Visual Odometry
Sources
arXiv 2108.10869: DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
princeton-vl/DROID-SLAM (GitHub)

See it in the full glossary →