Embodied AI Glossary中文

Depth Estimation

深度估计Common

Inferring how far every pixel in an image is from the camera, producing a depth map.

Depth estimation is the task of inferring, from an image, the distance from the camera to each point in the scene, usually producing a depth map the same size as the image. Roughly three approaches exist: binocular stereo matching, which computes distance by triangulation from the disparity between two cameras; active ranging, such as structured light, time-of-flight depth cameras, and lidar; and monocular depth estimation, where a neural network infers depth from a single image. Monocular estimation has an inherent scale ambiguity, so early models mostly gave only relative depth; MiDaS (TPAMI 2020) improved cross-scene generalization by training on a mix of multiple datasets, and more recent models such as Depth Pro and Metric3D can output metric depth in meters directly. For robots, depth is the foundation for turning pixels into point clouds and for grasping and obstacle avoidance; learned methods can also patch the holes that depth cameras leave on transparent or reflective objects.

ExamplePhotographing a glass cup with a depth camera leaves a large chunk of the cup's body with missing depth. Using a learned model to estimate depth from the RGB image instead fills in that hole, producing a complete point cloud usable for grasp detection.

Also called
Depth Prediction
Related
Monocular Depth Estimation · Stereo Matching · Depth Camera · Metric Depth / Relative Depth · Depth Completion · Depth Anything
Sources
Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer (MiDaS, arXiv 1907.01341)
Depth Anything V2 (arXiv 2406.09414)

See it in the full glossary →