Monocular Depth Estimation
单目深度估计MDECommonPredicting how far away every pixel is from just one ordinary color photo.
Monocular depth estimation predicts the depth of every pixel from just a single ordinary RGB image, taken with one camera rather than a stereo pair, producing a depth map. It's inherently ambiguous: the same photo could show a small nearby object or a large distant one, and scale is the hardest thing to pin down, so a model can only rely on learned cues such as typical object size, perspective, and occlusion. Researchers at NYU, David Eigen and colleagues, were among the first to use a deep network to predict depth coarse-to-fine in 2014; Depth Anything V2 (NeurIPS 2024) trains on synthetic data plus large-scale pseudo-labeled real images. Output is either relative depth (only the ordering of distances is known) or metric depth (given in meters). For robots, it can supply 3D information when no depth camera is available, and it's also commonly used to recover scene geometry from ordinary internet video.
ExampleAn arm has only a wrist-mounted RGB camera. Using the metric-depth version of Depth Anything V2 estimates a depth map from each frame, which is then back-projected into a point cloud using the camera's intrinsics for grasp planning.
- Also called
- MDE, Single-Image Depth Estimation
- Related
- Depth Estimation · Metric Depth / Relative Depth · Depth Anything · Depth Pro · Depth Map · Stereo Matching
- Sources
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network (arXiv:1406.2283)
Depth Anything V2 (arXiv:2406.09414)