3D Point Tracking
3D 点跟踪AdvancedContinuously tracks arbitrary pixels through a video and outputs their motion trajectory in 3D space.
3D point tracking is the three-dimensional version of tracking any point — following any chosen pixel through a video: given a video and a set of query points, it outputs each point’s 3D coordinates and whether it’s occluded at every frame. 2D tracking can’t tell whether the object is moving or the camera is, and it carries no depth; a 3D trajectory directly describes how an object moves and rotates in space. SpatialTracker (CVPR 2024) uses monocular depth estimation to lift pixels into 3D before tracking them; the follow-up SpatialTrackerV2 (ICCV 2025) folds point tracking, monocular depth, and camera pose estimation into one feed-forward model that works from monocular video alone. The TAPVid-3D benchmark is used to evaluate this task. In robot learning, 3D point trajectories can serve as an intermediate representation for turning human videos into robot actions — for example, General Flow predicts the future 3D trajectories of points on an object, guided by language instructions, to direct manipulation.
ExampleFilming someone pulling open a drawer and running SpatialTrackerV2 to track points on the handle gives a 3D trajectory moving outward in a straight line; from that, the drawer’s sliding direction can be inferred, and a robot can pull along the same direction.
- Also called
- TAP-3D, Tracking Any Point in 3D, SpatialTrackerV2
- Related
- Tracking Any Point · CoTracker · Scene Flow · Monocular Depth Estimation · 4D Reconstruction
- Sources
- SpatialTrackerV2: 3D Point Tracking Made Easy (arXiv 2507.12462)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D (arXiv 2407.05921)
General Flow as Foundation Affordance for Scalable Robot Learning (arXiv 2401.11439) - As of
- 2025-10