Embodied AI Glossary中文

Camera Depth Models

CDM 相机深度模型CDMAdvanced

A depth-repair model from ByteDance Seed that cleans up noisy depth-camera output into near-simulation-quality accurate depth.

CDM comes from the September 2025 paper “Manipulation as in Simulation,” released by ByteDance Seed together with Shanghai Jiao Tong University, Zhejiang University, and Tsinghua University. Ordinary depth cameras produce noisy, incomplete depth on reflective, transparent, thin, and edge regions, which makes it hard for policies trained on depth or point clouds to transfer directly from simulation to a real robot. CDM is a software plug-in that sits right after the camera: it takes an RGB image and raw depth as input and outputs denoised, metric depth (real-world scale, in meters). Its training data comes from the authors’ “neural data engine,” which generates paired data at scale by simulating the depth noise patterns of specific camera models; models are provided per camera model, covering several RealSense models, the Azure Kinect, and the ZED 2i, and both the models and the ByteCameraDepth dataset are open source. The paper shows that a manipulation policy trained only on clean simulated depth, once paired with CDM, can handle articulated, reflective, and thin objects on a real robot with no added noise and no fine-tuning.

ExampleA “put the bowl in the microwave” policy is trained in simulation using only clean depth images; at deployment, the RealSense’s raw depth is first passed through the matching CDM model, and the cleaned-up depth is then fed to the policy, with no real-robot data needed for fine-tuning.

Also called
CDM, Manipulation as in Simulation
Related
Depth Completion · Sim-to-Real Transfer · Sim-to-Real Gap (Reality Gap) · Depth Camera · Transparent & Reflective Object Perception · LingBot-Depth
Sources
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots (arXiv 2509.02530)
Manipulation as in Simulation 项目页 (Chinese)
As of
2025-09

See it in the full glossary →