RynnVLA-002
达摩院 RynnVLA-002AdvancedAn open-source VLA from Alibaba DAMO Academy that merges an action model and a world model into one autoregressive network.
RynnVLA-002 was released and open-sourced by Alibaba's DAMO Academy in November 2025. Its predecessor, RynnVLA-001 (August 2025), was a 7B-parameter VLA that first went through video-generation pretraining on first-person human videos, then transferred the manipulation skills it learned onto a robot arm. RynnVLA-002 merges two kinds of models into one: the VLA part outputs actions from images and language instructions, and the world-model part predicts the next frame from the current image and an action. Both share a single autoregressive backbone based on Chameleon, treating images, text, and actions all as tokens, plus a separate Action Transformer that outputs continuous actions, with support for wrist-camera and robot-state input. The authors' view is that learning to predict 'how the image changes as a result of an action' helps the model understand physical dynamics, which in turn improves action quality. The paper reports 97.4% success on the LIBERO simulation benchmark without pretraining, and roughly a 50% success-rate gain from adding the world model on real-robot LeRobot tasks.
ExampleThe same model can output the robot arm's next segment of motion for 'put the block in the box,' and can also generate the wrist-camera image that would result from executing a given action.
- Also called
- RynnVLA, RynnVLA-002: A Unified Vision-Language-Action and World Model
- Related
- World Action Model · World Model · Vision-Language-Action Model · WorldVLA · LIBERO Benchmark · RynnBrain
- Sources
- RynnVLA-002: A Unified Vision-Language-Action and World Model (arXiv 2511.17502)
alibaba-damo-academy/RynnVLA-002 (GitHub)
alibaba-damo-academy/RynnVLA-001 (GitHub) - As of
- 2026-05