Reward Engineering
奖励工程CommonDesigning and debugging a reward function for reinforcement learning so the robot actually learns the intended behavior.
Reward engineering covers the whole job of designing a reward function for a reinforcement-learning task: which terms to include, how heavily to weight each one, when to award them, and how to prevent the agent from gaming them. With only a sparse reward like “+1 on success,” the agent rarely stumbles onto positive feedback by chance, so dense intermediate rewards are often added instead — reward shaping — such as scoring higher the closer the agent gets to the goal. A robot locomotion reward often has a dozen or more terms: rewarding tracking of a commanded velocity while penalizing excessive torque, jerky motion, and body collisions, with weights that need extensive trial and error — get it wrong and reward hacking follows. Recent work also has large models write rewards automatically: Eureka has GPT-4 iteratively generate reward code, beating reward functions written by human experts on 83% of 29 tasks.
Examplelegged_gym's default quadruped-walking reward: tracking linear-velocity commands (weight 1.0), tracking angular velocity (0.5), penalizing vertical velocity (−2.0), penalizing joint torque and acceleration, penalizing the rate of change of actions (−0.01) and body collisions (−1), and rewarding time each foot spends airborne (1.0).
- Also called
- Reward Design
- Related
- Reward Function · Reward Shaping · Sparse Reward · Dense Reward · Reward Hacking · Eureka
- Sources
- Eureka: Human-Level Reward Design via Coding Large Language Models
legged_gym: legged_robot_config.py (reward scales)
Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning