Embodied AI Glossary中文

MegaPose

Advanced

A method that estimates 6D pose for a new object given only its CAD model, with no retraining needed.

MegaPose was proposed by Yann Labbé, Dieter Fox, Josef Sivic, and colleagues at Inria, NVIDIA, and other institutions, published at CoRL 2022. 6D pose estimation computes an object’s 3D position and orientation in the camera’s coordinate frame; earlier methods mostly needed to be trained separately for each object, so a new part meant starting over. MegaPose takes a “render and compare” approach: given the region of the image containing the object and its CAD model, it first roughly estimates a pose, renders a synthetic image of the model at that pose, compares it against the real image, and has a network predict how to correct the pose, repeating this iteratively. It is trained on large-scale, photorealistic synthetic data covering thousands of different objects; the paper argues that this object diversity is the key to generalization, letting the method work directly on hundreds of new objects with no retraining, achieving competitive results on benchmarks like YCB-Video and BOP. NVIDIA’s later FoundationPose used a similar render-and-compare refinement step and compared itself against MegaPose.

ExampleA factory receives a new type of part; given only its CAD file, a detector first boxes the part, then MegaPose estimates its 6D pose for a robot arm to grasp, with no need to collect and train on data specific to this part.

Also called
MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare
Related
6D Object Pose Estimation · FoundationPose · SAM-6D · BOP (Benchmark for 6D Object Pose Estimation) · YCB Object and Model Set · Pose Tracking
Sources
arXiv: MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

See it in the full glossary →