Embodied AI Glossary中文

CoTracker

Advanced

Meta’s open-source video point-tracking model that jointly tracks large numbers of pixels, including ones that get occluded.

CoTracker was proposed by Karaev and colleagues at Meta AI and Oxford’s VGG group, posted to arXiv in July 2023 and published at ECCV 2024. It belongs to the tracking-any-point family: given any pixel in a video, it outputs that point’s position and visibility in every subsequent frame. Most earlier methods tracked each point independently; CoTracker uses a Transformer to track a large batch of points jointly, exploiting the correlations between points to be more robust to occlusion and points leaving the frame, and it processes video in a sliding short time window so it can run online. The October 2024 CoTracker3 simplifies the architecture and trains using pseudo-labels generated on unlabeled real videos by existing models, needing roughly 1,000 times less training data than earlier methods. In robot learning, it’s commonly used to automatically label point trajectories on demonstration videos.

ExampleATM (Any-Point Trajectory Modeling) uses CoTracker to generate point trajectories on demonstration videos as ground truth to train a trajectory-prediction model, then uses the predicted future trajectories to guide a robot policy.

Also called
CoTracker: It is Better to Track Together, CoTracker3, CoTracker2
Related
Tracking Any Point · TAPIR · Optical Flow · ATM · 3D Point Tracking · Occlusion
Sources
CoTracker: It is Better to Track Together (arXiv 2307.07635)
CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos (arXiv 2410.11831)
facebookresearch/co-tracker (GitHub)
As of
2025-01

See it in the full glossary →