Embodied AI Glossary中文

Embodied Visual Tracking

具身视觉跟踪(目标跟随)EVTAdvanced

A robot using its own camera to keep following a specified target, keeping it in view continuously as it moves.

Embodied visual tracking means an agent, using only first-person vision, continuously follows a specified target, usually a pedestrian or another robot, in a dynamic environment, controlling its own motion so the target stays in view at an appropriate distance. It differs from traditional visual tracking in that traditional methods only draw a box around a target within pre-recorded video, whereas here the tracker must output its own motion commands, and what it sees next depends on how it chooses to move. Wenhan Luo and colleagues' 2018 active object tracking (ICML) used reinforcement learning to map images directly to actions such as moving forward and turning, an early landmark. The difficulty lies in occlusion, distractors that look similar to the target, sudden turns by the target, and having to do target recognition and path planning at the same time. TrackVLA (CoRL 2025), from teams at Peking University, Galbot, and others, uses a single vision-language-action model to handle both recognition and trajectory planning at once, and built the EVT-Bench benchmark, collecting about 1.7 million samples. Applications include companion following, camera-operator following, and logistics vehicles following a lead vehicle.

ExampleA robot dog told to “follow the person in red” locks onto that person in a crowd; when they turn into a hallway and are briefly blocked from view, the robot still finds them again and keeps following.

Also called
EVT, Active Object Tracking
Related
Object Tracking · TrackVLA · Vision-and-Language Navigation · Active Perception · Social Navigation · Vision-Language-Action Model
Sources
End-to-end Active Object Tracking via Reinforcement Learning (Luo et al., ICML 2018)
TrackVLA: Embodied Visual Tracking in the Wild
TrackVLA project page
As of
2025-05

See it in the full glossary →