Depth Anything
CommonA monocular depth estimation foundation model from HKU and TikTok that produces a depth map from a single photo.
Depth Anything is a series of monocular depth estimation models from research teams at the University of Hong Kong and TikTok. V1 (CVPR 2024) was trained on about 1.5 million labeled images plus about 62 million unlabeled images via pseudo-labeling, giving it strong generalization, and it outputs relative depth. V2 (NeurIPS 2024) instead trains a large teacher model on synthetic data and then trains a student model on large-scale pseudo-labeled real images, producing sharper detail; it comes in several sizes from about 25 million to 1.3 billion parameters, plus a metric-depth fine-tuned version. The Small version is licensed under Apache 2.0, while the larger versions are restricted to non-commercial use. In robotics it's often used to supplement depth cameras, providing depth priors for 3D perception, navigation, and data generation. A successor, Depth Anything 3, has since been released.
ExampleRunning Depth Anything V2-Small on a wrist camera's RGB frame gives per-pixel relative depth; fitting a scale and offset to a small number of points actually measured by a depth camera then recovers metric depth in meters.
- Also called
- Depth Anything V1, Depth Anything V2
- Related
- Monocular Depth Estimation · Metric Depth / Relative Depth · Depth Estimation · Depth Anything 3 · Prompt Depth Anything · Vision Foundation Model
- Sources
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data (arXiv 2401.10891)
Depth Anything V2 (arXiv 2406.09414)
DepthAnything/Depth-Anything-V2(GitHub) - As of
- 2025-01