Embodied AI Glossary中文

Residual Reinforcement Learning

残差强化学习Residual RLAdvanced

Keeping an existing base controller and using reinforcement learning to learn only a correction on top of its output.

Residual reinforcement learning was proposed in two nearly simultaneous papers at the end of 2018: UC Berkeley's Levine group working with Siemens on “Residual Reinforcement Learning for Robot Control,” and MIT's “Residual Policy Learning.” The final action equals the base policy's action plus a correction output by a residual policy; the base policy can be a hand-written controller, model-predictive control, or a policy learned through imitation, and reinforcement learning only has to fill in hard-to-model parts such as friction and contact. Because it starts from a point that already roughly works, exploration is safer and more sample-efficient, making it well suited to training directly on a real robot. In recent years it is often used to refine behavior-cloning policies: train and freeze a diffusion or action-chunking policy on demonstrations, then train a small closed-loop residual policy on top of it for real-time correction.

ExampleMIT's ResiP freezes a demonstration-trained action-chunking policy and uses it as a trajectory planner, then trains a closed-loop residual policy with reinforcement learning for real-time correction, applied to precision-assembly tasks where adding more behavior-cloning data no longer raises success rate.

Also called
Residual Policy Learning, RPL, Residual RL
Related
Residual Policy · Behavior Cloning · Real-World Reinforcement Learning · Reinforcement Fine-Tuning (RL Fine-Tuning) · Action Chunking · Noise-Space Policy Steering
Sources
Johannink et al. 2018: Residual Reinforcement Learning for Robot Control
Silver et al. 2018: Residual Policy Learning
Ankile et al. 2024: From Imitation to Refinement -- Residual RL for Precise Assembly

See it in the full glossary →