Embodied AI Glossary中文

Stereo Matching

立体匹配Advanced

Finding the same point’s position in both the left and right images to compute disparity, then converting that into depth.

Stereo matching is the core step of stereo depth sensing: the left and right images are first rectified onto the same horizontal lines (epipolar rectification), and then, for each pixel in the left image, the most similar point along the same row of the right image is found; the horizontal distance between them is the disparity, which is converted to depth via depth = focal length × baseline / disparity. Classic methods include block matching and semi-global matching (SGM); deep-learning methods build a cost volume with a neural network or refine it iteratively, as in RAFT-Stereo, and recent models such as FoundationStereo can generalize zero-shot to new scenes. The main difficulties are texture-less regions, reflective and transparent surfaces, and occlusion boundaries. Stereo matching determines the quality of a stereo camera’s depth map, and it’s often the part that gets replaced or improved when a robot needs to grasp transparent objects.

ExampleFeeding the left and right infrared images from a ZED or RealSense camera into FoundationStereo produces a more complete depth map than the camera’s own built-in algorithm.

Also called
Stereo Depth Estimation
Related
Disparity · Stereo Baseline · Stereo Camera · Epipolar Geometry · FoundationStereo · Depth Estimation
Sources
Middlebury Stereo Vision Page

See it in the full glossary →