VLA-Adapter
AdvancedA VLA that reaches top-tier performance without any robot pretraining, using a 0.5B small model plus a lightweight policy module.
VLA-Adapter was released in September 2025 by Beijing University of Posts and Telecommunications, Westlake University, Zhejiang University, HKUST (Guangzhou), the OpenHelix team, and others, with code and weights open-sourced. Typical VLAs rely on large vision-language models (VLMs) and pretraining on massive robot datasets, which is expensive. The authors systematically compared which layers and features of a VLM are best suited to condition action generation, and designed a roughly 97-million-parameter policy module accordingly: Bridge Attention injects raw vision-language features from various VLM layers, along with features from a set of learnable queries (ActionQuery), into the action space, with the amount injected controlled by learnable parameters. Using only Qwen2.5-0.5B as the backbone and no robot-data pretraining, it reaches 97.3% average success on LIBERO (98.5% for the Pro version) and an average completed length of 4.50 on CALVIN ABC→D (Pro version). This sharply lowers the barrier to training and deploying a VLA, and later work such as VLA-RFT builds on it.
ExampleThe paper reports a usable model can be trained on a single consumer GPU in about 8 hours; inference throughput reaches 219.2 Hz, compared with 71.4 Hz for OpenVLA-OFT under the same conditions.
- Also called
- VLA-Adapter-Pro, VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Related
- Vision-Language-Action Model · Learnable Query · Adapter · OpenVLA-OFT · LIBERO Benchmark · VLA-RFT
- Sources
- VLA-Adapter (arXiv 2509.09372)
VLA-Adapter 项目主页 (Chinese)
OpenHelix-Team/VLA-Adapter (GitHub) - As of
- 2025-09