Embodied AI Glossary中文

TAPIR

Advanced

A DeepMind point-tracking model that first finds a rough match per frame, then refines the trajectory over time.

TAPIR is a point-tracking model proposed in 2023 (ICCV 2023) by Google DeepMind and Oxford’s VGG group, for the “Tracking Any Point” (TAP) task: given any point on any frame of a video, output that point’s position in every other frame, and whether it is occluded. The method has two stages: a matching stage that independently finds, on every frame, the candidate location most similar to the query point, used as an initialization; and a refinement stage that uses local correlation to repeatedly update the whole trajectory and the query feature over time. The paper reports a clear improvement over prior methods on the TAP-Vid benchmark. TAPNet is the earlier baseline the same team introduced in the TAP-Vid paper, and its code repository still carries the name “tapnet.” In robotics, TAPIR has been used to track keypoints on objects or on a gripper — for instance, DeepMind’s RoboTAP uses point trajectories for few-shot imitation.

ExampleClicking on one corner of a towel in a video of a robot arm folding it, TAPIR outputs that corner’s pixel coordinates in every later frame, picking it back up even after the gripper briefly hides it.

Also called
Tracking Any Point with per-frame Initialization and temporal Refinement, TAPNet
Related
Tracking Any Point · CoTracker · Optical Flow · ATM · Keypoint Detection · Occlusion
Sources
TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement (arXiv 2306.08637)
As of
2023-06

See it in the full glossary →