Embodied AI Glossary中文

Visual Chain-of-Thought

视觉思维链Visual CoTAdvanced

Having a model first produce visual intermediate steps, like generated images or marked regions, before giving its final answer or action.

Chain-of-thought (CoT) originally meant having a large language model write out its reasoning steps before answering; visual chain-of-thought replaces those intermediate steps with visual ones. In multimodal models, the Visual CoT dataset released by Shao and colleagues in 2024 contains 438,000 question-answer pairs with intermediate bounding boxes, training models to first circle the key region of an image, look closer, and only then answer. In robotics, a more common pattern is to first “imagine” the goal image: CoT-VLA (CVPR 2025), from NVIDIA, Stanford, and others, first autoregressively generates an image of a future subgoal, then generates the action chunk to reach it, beating the previous best VLA by 17% on real robots and 6% in simulation. The benefit is that the target state is drawn out explicitly, which makes it easier to check and helps with long-horizon tasks; the cost is slower inference, since an extra image must be generated. It sits alongside other intermediate representations, such as embodied chain-of-thought, which writes out subtasks and object locations in text, and action chain-of-thought.

ExampleAfter receiving a manipulation instruction, CoT-VLA first generates an image of what the scene should look like several steps ahead — say, the object already grasped and near its target — as a subgoal, then generates a sequence of actions conditioned on that image. After execution, it observes again and imagines the next subgoal image.

Also called
Visual CoT, Visual Reasoning Chain
Related
Chain-of-Thought · Embodied Chain-of-Thought · Action Chain-of-Thought · CoT-VLA · Vision-Language-Action Model · Video Prediction Model
Sources
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models (arXiv:2503.22020)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning (arXiv:2403.16999)

See it in the full glossary →