Visuomotor Policy
视觉运动策略CommonA control policy that maps camera images directly to robot actions, usually a neural network.
A “policy” is a mapping from observations to actions; a visuomotor policy specifically takes input that is mostly images, often with proprioceptive state such as joint angles added, and outputs motor commands, joint positions, or end-effector pose. The term was popularized by Sergey Levine, Chelsea Finn, and colleagues' 2015 paper “End-to-End Training of Deep Visuomotor Policies”: using a convolutional network with about 92,000 parameters, it mapped raw images directly to joint torques to complete tasks such as screwing on a bottle cap, showing that training perception and control jointly worked better than training them separately. Today's ACT, Diffusion Policy, and language-augmented VLA models are all visuomotor policies, most commonly trained with imitation learning or reinforcement learning.
ExampleThe Diffusion Policy paper is literally titled “Visuomotor Policy Learning via Action Diffusion”: it takes in camera images and outputs a sequence of robot-arm actions to complete manipulation tasks such as pushing a T-shaped block.
- Related
- Policy · End-to-End · Diffusion Policy · Imitation Learning · Vision-Language-Action Model · End-to-End Training of Deep Visuomotor Policies
- Sources
- End-to-End Training of Deep Visuomotor Policies (arXiv 1504.00702)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)