Embodied AI Glossary中文

Homography

单应性矩阵Advanced

A 3×3 matrix describing how pixels on the same plane correspond between two images.

A homography H is a 3×3 matrix; because uniform scaling doesn’t change the mapping, it has only 8 degrees of freedom. Under the pinhole camera model, when the same physical plane is photographed from two viewpoints, corresponding pixels in the two images satisfy x′ ∝ Hx (with x in homogeneous coordinates); when a camera only rotates about its optical center with no translation, any two images of an arbitrary scene also satisfy a homography relationship. Solving for H needs at least 4 point correspondences, commonly using the direct linear transform (DLT), with RANSAC used to reject mismatches. Homographies are used for image stitching, perspective correction, and turning an image of the ground or a tabletop into a top-down view; they are also used in camera calibration — Zhang Zhengyou’s calibration method first solves for the homography from a checkerboard plane to each image, then decomposes the camera intrinsics from those. When photographing a general 3D scene with a camera that also translates, the fundamental or essential matrix from epipolar geometry is needed instead.

ExampleIn a tabletop manipulation experiment, sticking markers at the table’s four corners and measuring their coordinates on the table lets you solve for the homography from image to tabletop plane, converting a detected object’s pixel position directly into x, y coordinates on the table.

Also called
Homography Matrix, H Matrix, Homographic Transform
Related
Pinhole Camera Model · Camera Calibration · Epipolar Geometry · Feature Matching · Random Sample Consensus · Bird’s-Eye View
Sources
Wikipedia: Homography (computer vision)
Zhang: A Flexible New Technique for Camera Calibration(IEEE TPAMI 2000)

See it in the full glossary →