Embodied AI Glossary中文

3D Point Tracking

3D 点跟踪Advanced

Continuously tracks arbitrary pixels through a video and outputs their motion trajectory in 3D space.

3D point tracking is the three-dimensional version of tracking any point — following any chosen pixel through a video: given a video and a set of query points, it outputs each point’s 3D coordinates and whether it’s occluded at every frame. 2D tracking can’t tell whether the object is moving or the camera is, and it carries no depth; a 3D trajectory directly describes how an object moves and rotates in space. SpatialTracker (CVPR 2024) uses monocular depth estimation to lift pixels into 3D before tracking them; the follow-up SpatialTrackerV2 (ICCV 2025) folds point tracking, monocular depth, and camera pose estimation into one feed-forward model that works from monocular video alone. The TAPVid-3D benchmark is used to evaluate this task. In robot learning, 3D point trajectories can serve as an intermediate representation for turning human videos into robot actions — for example, General Flow predicts the future 3D trajectories of points on an object, guided by language instructions, to direct manipulation.

ExampleFilming someone pulling open a drawer and running SpatialTrackerV2 to track points on the handle gives a 3D trajectory moving outward in a straight line; from that, the drawer’s sliding direction can be inferred, and a robot can pull along the same direction.

Also called
TAP-3D, Tracking Any Point in 3D, SpatialTrackerV2
Related
Tracking Any Point · CoTracker · Scene Flow · Monocular Depth Estimation · 4D Reconstruction
Sources
SpatialTrackerV2: 3D Point Tracking Made Easy (arXiv 2507.12462)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D (arXiv 2407.05921)
General Flow as Foundation Affordance for Scalable Robot Learning (arXiv 2401.11439)
As of
2025-10

See it in the full glossary →