Embodied AI Glossary中文

Visuomotor Policy

视觉运动策略Common

A control policy that maps camera images directly to robot actions, usually a neural network.

A “policy” is a mapping from observations to actions; a visuomotor policy specifically takes input that is mostly images, often with proprioceptive state such as joint angles added, and outputs motor commands, joint positions, or end-effector pose. The term was popularized by Sergey Levine, Chelsea Finn, and colleagues' 2015 paper “End-to-End Training of Deep Visuomotor Policies”: using a convolutional network with about 92,000 parameters, it mapped raw images directly to joint torques to complete tasks such as screwing on a bottle cap, showing that training perception and control jointly worked better than training them separately. Today's ACT, Diffusion Policy, and language-augmented VLA models are all visuomotor policies, most commonly trained with imitation learning or reinforcement learning.

ExampleThe Diffusion Policy paper is literally titled “Visuomotor Policy Learning via Action Diffusion”: it takes in camera images and outputs a sequence of robot-arm actions to complete manipulation tasks such as pushing a T-shaped block.

Related
Policy · End-to-End · Diffusion Policy · Imitation Learning · Vision-Language-Action Model · End-to-End Training of Deep Visuomotor Policies
Sources
End-to-End Training of Deep Visuomotor Policies (arXiv 1504.00702)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)

See it in the full glossary →