Embodied AI Glossary中文

RT-2

Essential

Google DeepMind's 2023 model that outputs robot actions as text tokens — the paper that coined the term VLA.

RT-2 is a model Google DeepMind released in July 2023, and its paper was the first to use the term “vision-language-action model” (VLA). The approach: take a vision-language model already pretrained on internet image-text data (PaLI-X or PaLM-E), discretize each dimension of a robot action into 256 bins, write them out as a string of number tokens, and have the model output them just like ordinary text. During training, RT-1's robot data is co-fine-tuned together with web-scale visual question-answering data, so the model doesn't lose its existing knowledge. This lets the robot draw directly on commonsense knowledge learned from the web, producing what the paper calls emergent capabilities: recognizing objects and symbols absent from the robot data, and understanding semantic concepts like “the smallest one” or “something that could be used as a hammer.” The model comes in two sizes, 12 billion parameters (PaLM-E version) and 55 billion parameters (PaLI-X version), and it opened the path that later VLA models such as OpenVLA and π0 followed.

ExampleTold to “move the banana to the sum of two plus one,” RT-2 can place the banana on the spot marked 3, even though the robot's own training data contains no such arithmetic task.

Also called
Robotics Transformer 2, RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Related
Vision-Language-Action Model · RT-1 · PaLM-E · Co-training · Action Binning · OpenVLA
Sources
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv 2307.15818)
RT-2 项目主页 (Chinese)
As of
2023-07

See it in the full glossary →