Embodied AI Glossary中文

DreamVLA

Advanced

A VLA model that first predicts future dynamic regions, depth, and semantic features, then generates actions conditioned on them.

DreamVLA is a vision-language-action (VLA) model proposed in July 2025 by researchers from Shanghai Jiao Tong University, the Eastern Institute of Technology in Ningbo, Tsinghua University, Peking University, Galbot, and other institutions, published at NeurIPS 2025. Some VLA models first “imagine” a complete next frame before producing an action, but most pixels in a full image are irrelevant to the task. DreamVLA instead predicts only three compact kinds of “world knowledge”: where things will move (dynamic regions), monocular depth, and high-level semantics (using features from DINOv2 and SAM); conditioned on these, a diffusion Transformer then generates a segment of future actions — effectively predicting the outcome first and working backward to figure out what to do, an inverse-dynamics-style approach. To keep the three kinds of information from interfering with each other inside attention, it separates them with a block-structured attention mask. It is a representative example of the “VLA combined with world-model-style prediction” direction.

ExampleOn the CALVIN ABC-D long-horizon benchmark, it completes an average of 4.44 consecutive sub-tasks, and reaches a 76.7% success rate on real-robot manipulation tasks.

Also called
A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
Related
Vision-Language-Action Model · World Model · Inverse Dynamics Model · Diffusion Transformer · Attention Mask · CALVIN Benchmark
Sources
DreamVLA (arXiv 2507.04447)
DreamVLA 项目主页 (Chinese)
DreamVLA 代码仓库(GitHub) (Chinese)
As of
2025-09

See it in the full glossary →