Safe Reinforcement Learning
安全强化学习Safe RLAdvancedMaximizing reward while guaranteeing safety constraints are not violated, during both training and deployment.
García and Fernández's 2015 survey defines safe reinforcement learning as maximizing expected return, during learning and/or deployment, in problems that require reasonable performance guarantees or must respect safety constraints. Two broad approaches are common. One changes the optimization objective, for example by formulating the problem as a constrained Markov decision process, which adds an upper limit on “cost” alongside reward, and solving it with Lagrangian multipliers or a method such as 2017's Constrained Policy Optimization (CPO). The other changes the exploration process itself, for example by injecting prior knowledge, or by using a safety filter or control barrier function to intercept dangerous actions before they execute. This matters enormously for real-robot RL and for humanoid or legged locomotion control, where falling, collisions, or exceeding a joint's limits can damage hardware or even injure people.
ExampleThe CaT method rewrites every constraint in a legged robot's locomotion control as “end the episode early, with some probability, the moment it's violated”; with only a small change to PPO, it learned to cross obstacles on a real Solo quadruped robot.
- Also called
- Safe RL, Constrained Reinforcement Learning, Constrained RL
- Related
- Embodied Safety · Control Barrier Function · Safety Filter · Real-World Reinforcement Learning · Proximal Policy Optimization · Reward Shaping
- Sources
- García & Fernández 2015: A Comprehensive Survey on Safe Reinforcement Learning (JMLR)
Achiam et al. 2017: Constrained Policy Optimization
Chane-Sane et al. 2024: CaT: Constraints as Terminations for Legged Locomotion Reinforcement Learning