Off-Policy
异策略CommonReinforcement learning that can use data collected by a different policy, not only data the current policy gathered itself.
Reinforcement learning distinguishes two roles: the behavior policy, which interacts with the environment and produces data, and the target policy, the one actually being learned and improved. When these can differ, it's called off-policy learning. This means experience from older versions of the policy, human demonstrations, or data from other algorithms can all be stored in an experience-replay buffer and reused repeatedly, giving good sample efficiency; Q-learning, DQN, DDPG, TD3, and SAC all belong to this family. The cost is that the data distribution doesn't match the current policy, which makes training more prone to instability. Because real-robot interaction is expensive, real-robot reinforcement learning often chooses off-policy algorithms. Note that although this idea is sometimes loosely referred to with the same Chinese phrase as “offline,” off-policy is not the same thing as offline reinforcement learning — an off-policy algorithm is usually still collecting new data while it trains.
ExampleSERL uses RLPD, an off-policy algorithm built on SAC, with each training batch drawn half from human demonstrations and half from replay data the robot collects online, training policies for tasks like PCB insertion and cable routing in 25 to 50 minutes on average.
- Also called
- Off-Policy Learning
- Related
- On-Policy · Experience Replay · Soft Actor-Critic · Q-Learning · Offline Reinforcement Learning · SERL
- Sources
- OpenAI Spinning Up: Kinds of RL Algorithms
Wikipedia: State–action–reward–state–action (SARSA vs Q-learning)
SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning