Embodied AI Glossary中文

LingBot-Depth

蚂蚁灵波 LingBot-DepthAdvanced

An open-source depth-completion model from Ant Group’s Robbyant that uses a color image to repair a depth camera’s holes and noise.

LingBot-Depth is a depth model released and open-sourced in January 2026 by Robbyant, the embodied-AI company under Ant Group; the paper is titled “Masked Depth Modeling for Spatial Perception,” and the GitHub page states it has been accepted at ECCV 2026. RGB-D depth cameras often can’t measure depth on transparent or reflective surfaces, leaving holes and noise. LingBot-Depth treats these missing regions as a natural “mask”: given a color image, raw depth, and camera intrinsics, it uses a ViT-Large backbone to fuse the two modalities and outputs a completed metric depth map and a point cloud in the camera’s coordinate frame. It was trained on about 3 million paired RGB-D samples (about 2 million real, 1 million synthetic). The team states its depth-completion error is 40–50% lower than the best prior methods. Code, weights, and the dataset are all open source.

ExampleIn the official grasping experiments, grasping hard-to-sense objects using depth repaired by LingBot-Depth: success rate for a transparent storage box rose from 0% to 50%, for a glass cup from 60% to 80%, and for a steel cup from 65% to 85%.

Also called
Masked Depth Modeling for Spatial Perception, Masked Depth Modeling
Related
Depth Completion · Depth Camera · Transparent & Reflective Object Perception · Depth Holes · Masked Autoencoder · Robbyant
Sources
Masked Depth Modeling for Spatial Perception (arXiv 2601.17895)
GitHub: Robbyant/lingbot-depth
Robbyant 官网:LingBot-Depth (Chinese)
As of
2026-09

See it in the full glossary →