UniVLA
AdvancedA framework that trains a cross-embodiment VLA by learning 'task-relevant latent actions' from video.
UniVLA was released by the University of Hong Kong's OpenDriveLab and AgiBot in May 2025, published at RSS 2025. Most VLAs (vision-language-action models) rely on large amounts of action-labeled robot data and are tied to a single robot. UniVLA first trains a latent action model: looking at two consecutive frames in DINOv2 feature space, together with the language instruction, it separates task-relevant changes from irrelevant ones like camera shake, and quantizes the relevant changes into discrete latent action tokens, which lets videos without action labels — including human videos — be used for pretraining too. It then uses Prismatic-7B as the backbone to predict these latent actions, and when deploying to a specific robot, only adds a small decoder head of about 12 million parameters to translate them into real actions. The paper reports that with under 1/20th of OpenVLA's pretraining compute and only 1/10th of its downstream data, it outperforms OpenVLA on benchmarks including LIBERO, CALVIN, and R2R.
ExampleRobot arm data, navigation data, and human manipulation videos are all fed together into UniVLA's latent action model; the same learned latent actions, paired with different small decoder heads, can then drive a robot arm in LIBERO simulation and a navigation agent in R2R.
- Also called
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- Related
- Latent Action Model · Latent Action · Vision-Language-Action Model · OpenVLA · LAPA · Cross-Embodiment
- Sources
- UniVLA (arXiv 2505.06111)
OpenDriveLab/UniVLA GitHub - As of
- 2025-05