Embodied AI Glossary中文

RT-Trajectory

Advanced

A policy that replaces language instructions with a trajectory sketch drawn on the image, showing the robot how the motion should go.

RT-Trajectory was released by Google DeepMind together with UC San Diego, Stanford, and Intrinsic in November 2023, selected as an ICLR 2024 Spotlight. A language instruction only says 'what to do,' which is often not specific enough for tasks the model never saw during training. RT-Trajectory instead conditions on a 'trajectory sketch': the path the end effector should follow is drawn as a curve overlaid on the camera image, with color encoding time and height, and marks showing where the gripper should open or close. No manual annotation is needed for training — the end-effector positions recorded in each demonstration are simply projected onto the image after the fact to generate the sketch (hindsight). The policy backbone reuses RT-1. At test time, sketches can be hand-drawn by a person, extracted from human videos, or generated by a large model. Across 7 tasks unseen during training, it clearly outperforms language-conditioned baselines like RT-1 and RT-2, as well as goal-image-conditioned baselines.

ExampleTraining data is mostly pick-and-place; at test time, a person draws a curve on the image that first grabs one corner of a cloth and then pulls it toward the opposite side, and the robot follows it to perform 'fold the cloth,' a task absent from training.

Also called
RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches
Related
RT-1 · RT-2 · Intermediate Representation · Task Generalization · Hindsight Relabeling · Goal-conditioned Policy
Sources
RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches (arXiv 2311.01977)
RT-Trajectory project page
As of
2024-01

See it in the full glossary →