Embodied AI Glossary中文

Pose Tracking

位姿跟踪Advanced

Continuously estimating an object’s 3D position and orientation across a video’s successive frames.

Pose means an object’s 3D position plus 3D orientation, 6 degrees of freedom in total, which is why it is also called 6D pose. 6D pose estimation usually computes the pose from scratch in a single frame; pose tracking instead uses the previous frame’s result and makes only a small correction in the new frame, which makes it faster and smoother, and suitable for real-time closed-loop control. NVIDIA’s FoundationPose (a CVPR 2024 Highlight paper) unifies estimation and tracking in one framework: given an RGB-D image, it needs only a CAD model or a handful of reference images for objects it has never seen, and refines the pose iteratively using a “render-and-compare” approach. Robots rely on it when grasping moving objects, adjusting an object in-hand, or doing visual servoing; if the object is occluded or moves too fast, tracking can be lost and a fresh global estimate is needed. Estimating a camera’s own pose in SLAM is also sometimes called tracking, but that is a different target.

ExampleWhile a robot unscrews a bottle cap, FoundationPose tracks the bottle’s 6D pose every frame, and the controller adjusts the gripper position accordingly.

Also called
6D Pose Tracking, Object Pose Tracking
Related
6D Object Pose Estimation · FoundationPose · Visual Servoing · In-hand Manipulation · Occlusion · Pose
Sources
FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects (arXiv 2312.08344)
FoundationPose 项目主页 (NVIDIA) (Chinese)
As of
2024-06

See it in the full glossary →