Embodied AI Glossary中文

VLA-Adapter

Advanced

A VLA that reaches top-tier performance without any robot pretraining, using a 0.5B small model plus a lightweight policy module.

VLA-Adapter was released in September 2025 by Beijing University of Posts and Telecommunications, Westlake University, Zhejiang University, HKUST (Guangzhou), the OpenHelix team, and others, with code and weights open-sourced. Typical VLAs rely on large vision-language models (VLMs) and pretraining on massive robot datasets, which is expensive. The authors systematically compared which layers and features of a VLM are best suited to condition action generation, and designed a roughly 97-million-parameter policy module accordingly: Bridge Attention injects raw vision-language features from various VLM layers, along with features from a set of learnable queries (ActionQuery), into the action space, with the amount injected controlled by learnable parameters. Using only Qwen2.5-0.5B as the backbone and no robot-data pretraining, it reaches 97.3% average success on LIBERO (98.5% for the Pro version) and an average completed length of 4.50 on CALVIN ABC→D (Pro version). This sharply lowers the barrier to training and deploying a VLA, and later work such as VLA-RFT builds on it.

ExampleThe paper reports a usable model can be trained on a single consumer GPU in about 8 hours; inference throughput reaches 219.2 Hz, compared with 71.4 Hz for OpenVLA-OFT under the same conditions.

Also called
VLA-Adapter-Pro, VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
Related
Vision-Language-Action Model · Learnable Query · Adapter · OpenVLA-OFT · LIBERO Benchmark · VLA-RFT
Sources
VLA-Adapter (arXiv 2509.09372)
VLA-Adapter 项目主页 (Chinese)
OpenHelix-Team/VLA-Adapter (GitHub)
As of
2025-09

See it in the full glossary →