Vision-Language-Action Model
视觉-语言-动作模型VLAEssentialA large model that looks at images, understands language instructions, and directly outputs robot actions.
A vision-language-action model takes camera images and a natural-language instruction as input and directly outputs robot control actions, usually built by taking a pretrained vision-language model, a large model that can look at images and answer questions, and training it further on robot data. The name comes from Google DeepMind's RT-2 paper in July 2023: it writes actions as text tokens and trains them together with web image-text data, letting the robot draw on internet knowledge to handle objects and instructions it has never seen. Since then there has been the open-source OpenVLA (7B parameters, trained on 970,000 robot demonstrations), and Physical Intelligence's π0, which generates continuous action chunks with flow matching. The main difference between them is how actions get output: discrete tokens, a diffusion or flow-matching action head, or some mix of the two.
ExampleGiven a tabletop photo and the instruction “put the eggplant in the pot,” OpenVLA outputs 7 action tokens, which get decoded into the arm end-effector's translation, rotation, and gripper open/close.
- Also called
- VLA, VLA Model
- Related
- Vision-Language Model · RT-2 · OpenVLA · π0 · Action Tokenizer · Action Expert
- Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv:2307.15818)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164) - As of
- 2024-10