Embodied AI Glossary中文

ATM

ATM(任意点轨迹建模)Advanced

Learns how any point in a scene will move from video, then uses that predicted motion to guide a robot policy.

ATM was proposed by researchers at UC Berkeley, Tsinghua's Institute for Interdisciplinary Information Sciences, and collaborators (including Yang Gao and Pieter Abbeel), published at RSS 2024. Action-labeled robot data is expensive; action-free video is abundant. Earlier video pretraining mostly predicted future frames pixel by pixel, which is computationally heavy and full of irrelevant detail. ATM instead predicts point motion: it first uses the point tracker CoTracker to label 2D trajectories for points in a video, then trains a Transformer to predict, from the current image, a language instruction, and a set of point positions, where those points will go in the future; a policy is then trained to output actions using the predicted trajectories as a subgoal, needing only a small number of action-labeled demonstrations. Across more than 130 tasks including LIBERO, it beat video-pretraining baselines by about 80% on average, and can also transfer skills from human videos.

ExampleGiven the instruction “open the middle drawer of the cabinet,” the model first draws, on the current image, how points such as the drawer handle will move as it's pulled open, and the policy then outputs arm actions guided by these predicted trajectories.

Also called
Any-point Trajectory Modeling, Any-point Trajectory Modeling for Policy Learning
Related
Tracking Any Point · CoTracker · Action-free Video · Pretraining on Human Videos · LIBERO Benchmark · Intermediate Representation
Sources
Any-point Trajectory Modeling for Policy Learning (arXiv 2401.00025)
ATM 项目页 (Chinese)
Robotics: Science and Systems XX (RSS 2024) 论文集 (Chinese)
As of
2024-07

See it in the full glossary →