RoboVLMs
AdvancedA systematic study and open-source framework comparing which VLM backbone, architecture, and training data work best for building a VLA.
RoboVLMs was released in December 2024 by Tsinghua University, ByteDance Research, the Institute of Automation at the Chinese Academy of Sciences, Shanghai Jiao Tong University, the National University of Singapore, and others. It is both a systematic experimental study and an open-source framework of the same name that makes it easy to plug a new vision-language model into a VLA. It answers several of the key design choices when building a VLA: which VLM backbone to use, how to organize action output and history information, and when to bring in cross-embodiment data. The authors compared more than 8 VLM backbones and 4 policy architectures across over 600 experiments. The main findings: an independent policy head that outputs continuous actions and takes in multiple frames of history performs best; backbones with thorough vision-language pretraining, such as KosMos and PaliGemma, do noticeably better; and pretraining on cross-embodiment data first helps robustness. The best configuration averaged 4.49 consecutive completed tasks on CALVIN.
ExampleTo try a new vision-language model as a VLA backbone, one can swap just the backbone inside the RoboVLMs framework while keeping everything else fixed, and compare it directly against the KosMos and PaliGemma versions on CALVIN.
- Also called
- Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Models, What Matters in Building Vision-Language-Action Models for Generalist Robots
- Related
- Vision-Language-Action Model · RoboFlamingo · Action Head · Cross-Embodiment Data · PaliGemma · Ablation Study
- Sources
- arXiv 2412.14058: RoboVLMs
RoboVLMs 项目主页 (Chinese) - As of
- 2026-02