Trust Region Policy Optimization
信赖域策略优化TRPOAdvancedA policy-gradient algorithm that bounds the KL divergence between old and new policy at every update; PPO's predecessor.
Trust region policy optimization was proposed by Schulman, Levine, Abbeel, and colleagues at ICML 2015. An ordinary policy-gradient step can, if too large, cause performance to collapse, and even a small change in parameters can shift the action distribution a great deal. TRPO instead puts the constraint on the policy's distribution: it maximizes a surrogate objective within a “trust region” where the average KL divergence between old and new policy stays under a threshold, which in theory guarantees monotonic improvement. In practice it uses conjugate gradient to find an approximate second-order update direction, then a backtracking line search to check the constraint and the improvement. It is an on-policy algorithm and fairly involved to implement; PPO, proposed by the same authors in 2017, approximates the same constraint far more simply and has become the mainstream choice for robot reinforcement learning.
ExampleIn OpenAI's Spinning Up implementation of TRPO, each iteration first computes an update direction with conjugate gradient, then keeps shrinking the step size by a backtracking factor until the new policy stays within the KL limit and improves the surrogate objective.
- Also called
- TRPO
- Related
- Proximal Policy Optimization · Policy Gradient · Kullback-Leibler Divergence · On-Policy · Generalized Advantage Estimation · Reinforcement Learning
- Sources
- Schulman et al. 2015: Trust Region Policy Optimization (ICML 2015)
OpenAI Spinning Up: Trust Region Policy Optimization
Schulman et al. 2017: Proximal Policy Optimization Algorithms