RT-2
EssentialGoogle DeepMind's 2023 model that outputs robot actions as text tokens — the paper that coined the term VLA.
RT-2 is a model Google DeepMind released in July 2023, and its paper was the first to use the term “vision-language-action model” (VLA). The approach: take a vision-language model already pretrained on internet image-text data (PaLI-X or PaLM-E), discretize each dimension of a robot action into 256 bins, write them out as a string of number tokens, and have the model output them just like ordinary text. During training, RT-1's robot data is co-fine-tuned together with web-scale visual question-answering data, so the model doesn't lose its existing knowledge. This lets the robot draw directly on commonsense knowledge learned from the web, producing what the paper calls emergent capabilities: recognizing objects and symbols absent from the robot data, and understanding semantic concepts like “the smallest one” or “something that could be used as a hammer.” The model comes in two sizes, 12 billion parameters (PaLM-E version) and 55 billion parameters (PaLI-X version), and it opened the path that later VLA models such as OpenVLA and π0 followed.
ExampleTold to “move the banana to the sum of two plus one,” RT-2 can place the banana on the spot marked 3, even though the robot's own training data contains no such arithmetic task.
- Also called
- Robotics Transformer 2, RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Related
- Vision-Language-Action Model · RT-1 · PaLM-E · Co-training · Action Binning · OpenVLA
- Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv 2307.15818)
RT-2 项目主页 (Chinese) - As of
- 2023-07