Embodied AI Glossary中文

RT-H

Advanced

A hierarchical VLA from Google that has the robot first state a 'language motion' like 'move arm forward' before outputting the action.

RT-H was released by Google DeepMind and Stanford University in March 2024, with authors including Suneel Belkhale and Dorsa Sadigh. VLAs like RT-2 map directly from a task instruction, such as 'put the soda can in the drawer,' to motor commands, which makes it hard to learn the action structure shared across different tasks. RT-H inserts a middle layer of 'language motions' — fine-grained phrases like 'move arm forward' or 'close gripper': the same vision-language model, co-trained with internet data, first predicts a language motion from the task and the image, then outputs the specific action based on that language motion and the image. This lets semantically different tasks share the same underlying motions, and a person can correct the robot mid-execution simply by speaking, with those correction episodes then reused for further training. The paper reports roughly 15% better performance than RT-2 on multi-task data, and that learning from language interventions works better than learning from teleoperated interventions.

ExampleIf the robot's hand drifts off-target while opening a drawer, a person can just say 'move your arm to the left'; RT-H treats this as a new language motion and continues execution, and the correction is logged for further training.

Also called
RT-Hierarchy, RT-H: Action Hierarchies Using Language
Related
RT-2 · Hierarchical Architecture · Language Corrections · Human-in-the-Loop · Intermediate Representation · Vision-Language-Action Model
Sources
RT-H: Action Hierarchies Using Language (arXiv 2403.01823)
RT-H project page
As of
2024-06

See it in the full glossary →