CoT-VLA
AdvancedA 7-billion-parameter VLA that first generates a future subgoal image as its “thought,” then outputs an action to reach it.
CoT-VLA was proposed in March 2025 by NVIDIA, Stanford, MIT, and others, published at CVPR 2025. Earlier VLAs mapped images and instructions directly to actions with no reasoning step in between. CoT-VLA replaces chain-of-thought with a visual form: the model first autoregressively generates a subgoal frame several steps into the future — what the scene should look like once part of the task is done — then generates an action chunk targeting that image, executing it in closed loop. Its backbone is VILA-U, a multimodal model that can both understand and generate images, at 7B parameters total; image generation uses causal attention, while action decoding uses full attention. Because predicting a subgoal image needs no action labels, action-label-free video like EPIC-KITCHENS can be used for training too. The paper reports beating the strongest VLA baseline of the time by 17% on real-robot tasks and 6% on simulation benchmarks.
ExampleGiven “put the bowl in the drawer,” CoT-VLA first draws what the scene should look like a few steps ahead — the bowl already picked up and near the drawer — then outputs a chunk of arm actions to reach that image; after executing, it looks at the new image and repeats.
- Also called
- CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Visual Chain-of-Thought VLA
- Related
- Visual Chain-of-Thought · Vision-Language-Action Model · Chain-of-Thought · Action Chunking · Action-free Video · SuSIE
- Sources
- CoT-VLA (arXiv:2503.22020, CVPR 2025)
CoT-VLA 项目主页 (Chinese) - As of
- 2025-03