Embodied AI Glossary中文

Disparity

视差Advanced

The horizontal position difference of the same point between the left and right images — bigger for closer objects.

Disparity is the fundamental quantity in binocular stereo vision: with two cameras mounted side by side and the images rectified (so the same point falls on the same row in both), the difference d = xL − xR between a point’s x-coordinate in the left image and its x-coordinate in the right image is its disparity. Disparity is inversely proportional to depth: Z = f × B / d, where f is the focal length in pixels and B is the baseline distance between the two cameras. Computing disparity pixel by pixel gives a disparity map, which converts to a depth map with this formula. The process of computing disparity is called stereo matching; classic methods include OpenCV’s StereoBM and SGBM, and recent deep-learning methods include RAFT-Stereo and FoundationStereo. Because of the inverse relationship, distant objects have a disparity of only a few pixels or less, so a small matching error causes a large depth error — stereo ranging accuracy drops off quickly with distance, and lengthening the baseline improves accuracy at range.

ExampleWith focal length f = 640 pixels and baseline B = 5 centimeters, a point with disparity 16 pixels has depth Z = 640 × 0.05 / 16 = 2 meters; if the disparity is off by 1 pixel to 15, the computed depth becomes about 2.13 meters.

Also called
Disparity Map, Stereo Disparity
Related
Stereo Camera · Stereo Matching · Stereo Baseline · Depth Map · Epipolar Geometry · FoundationStereo
Sources
Wikipedia: Computer stereo vision

See it in the full glossary →