Embodied AI Glossary中文

Action Chain-of-Thought

动作思维链ACoTAdvanced

Having a VLA first reason out a coarse trajectory in action space, then generate the fine-grained action from it.

Chain-of-thought originally meant having a large model write out intermediate reasoning steps before giving an answer. Existing reasoning approaches in VLAs mostly predict subtask text or generate a goal image as the intermediate step — neither of which is an action itself. In January 2026, a team from Beihang University and AgiBot proposed ACoT-VLA (accepted to CVPR 2026), arguing for reasoning directly in action space: an Explicit Action Reasoner (EAR) uses flow matching to first generate a coarse reference trajectory; an Implicit Action Reasoner (IAR) uses learnable queries to extract a latent action prior from the VLM's internal features; the two are then combined via cross-attention to guide the action head in producing the final action sequence. The paper reports a 98.5% average success rate on LIBERO. It can be seen as another form of reasoning alongside embodied chain-of-thought and visual chain-of-thought. Separately, some 2026 navigation work also uses 'Action-CoT' to refer to step-by-step action reasoning.

ExampleACoT-VLA reports a 66.7% average success rate across three manipulation tasks on a real AgiBot G1 robot, and an 88.0% average success rate on LIBERO-Plus after supervised fine-tuning.

Also called
ACoT, ACoT-VLA
Related
Chain-of-Thought · Embodied Chain-of-Thought · Visual Chain-of-Thought · Vision-Language-Action Model · Action Head · Flow Matching
Sources
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models (arXiv 2601.11404)
ACoT-VLA 论文 HTML 版 (Chinese)
As of
2026-01

See it in the full glossary →