Embodied Chain-of-Thought
具身思维链ECoTCommonA method that has a robot model reason step by step — plan, subtask, object locations — before producing an action.
Embodied Chain-of-Thought (ECoT) was proposed in 2024 by researchers from UC Berkeley, Stanford, and other institutions. A standard vision-language-action (VLA) model looks at an image, hears an instruction, and outputs an action directly. ECoT instead has the model write out a chain of reasoning first — a restatement of the task, an overall plan, the current subtask, the direction the gripper should move next, the gripper's position, and bounding boxes for objects in the scene — and only then output the action. Unlike chain-of-thought in large language models, this reasoning has to stay grounded in the actual image and robot state. The authors used off-the-shelf foundation models to automatically label the BridgeData V2 dataset with these reasoning traces, then trained OpenVLA on them; on hard generalization tasks, absolute success rate rose by 28%, with no extra robot data. The cost is generating many more tokens, which slows inference down; a 2025 follow-up from the same team introduced a lighter training scheme that runs about 3x faster.
ExampleGiven the instruction ‘put the mushroom in the pot,’ an ECoT model first writes out a plan (find the mushroom → pick it up → move it above the pot → put it down), the current subtask (‘pick up the mushroom’), a movement direction (‘down and to the left’), and bounding boxes for the mushroom, the pot, and the gripper — only then does it output the 7-dimensional arm action.
- Also called
- ECoT, Embodied Chain-of-Thought Reasoning
- Related
- Chain-of-Thought · Embodied Reasoning · Vision-Language-Action Model · OpenVLA · Visual Chain-of-Thought · Inference Latency
- Sources
- Robotic Control via Embodied Chain-of-Thought Reasoning (arXiv 2407.08693)
Embodied Chain-of-Thought Reasoning 项目页 (Chinese)
Training Strategies for Efficient Embodied Reasoning (arXiv 2505.08243) - As of
- 2025-05