WorldVLA
AdvancedAlibaba DAMO Academy's model that merges a VLA and a world model into one autoregressive system, outputting actions and predicting the next frame.
WorldVLA was released by Alibaba's DAMO Academy in June 2025, built on an autoregressive image-and-text model in the Chameleon family, converting images, text, and actions all into tokens fed into the same Transformer. The single model plays two roles at once: as an action model, it generates actions from an image and an instruction; as a world model, it predicts the next frame from the current image and an action. The authors found that training the two jointly improves both. They also found that when actions are generated autoregressively one after another, errors in earlier actions propagate to later ones, so they introduce an attention mask that, when generating the current action, hides previous actions and looks only at the image and instruction, which noticeably improves multi-step action generation on LIBERO. The project was upgraded to RynnVLA-002 in November 2025.
ExampleOn a LIBERO simulation task, the same model can both output the robot arm's next segment of action and generate the resulting frame given a specified action.
- Also called
- WorldVLA: Towards Autoregressive Action World Model
- Related
- World Action Model · Vision-Language-Action Model · Autoregressive Decoding · Attention Mask · RynnVLA-002 · Alibaba DAMO Academy
- Sources
- WorldVLA (arXiv:2506.21539)
alibaba-damo-academy/WorldVLA (GitHub) - As of
- 2025-11