Embodied AI Glossary中文

Dense Reward

稠密奖励Common

A reward design that gives informative feedback at nearly every step, showing whether the agent is getting closer to or further from the goal.

Dense reward means a reinforcement-learning environment gives an informative reward at almost every timestep, as opposed to sparse reward, which pays off only on task success. It's usually built by a human, breaking the task into several weighted terms added together — distance to the goal, velocity-tracking error, an energy penalty, and so on. The benefit is that the agent knows which direction to improve in from the very start, so learning is faster; legged-locomotion reinforcement learning almost always uses dense reward. The cost is that it's laborious to design, and if the weights aren't tuned well, the agent will find a shortcut that racks up score without actually accomplishing the real goal — reward hacking. Projects like Eureka try to have a large language model write this kind of reward-function code automatically.

ExampleGymnasium-Robotics' FetchReach task has two versions: the sparse version gives −1 per step while the end effector is more than 5 cm from the target and 0 on arrival, while the dense version, FetchReachDense, gives the negative Euclidean distance to the target at every step. legged_gym's quadruped-walking reward likewise combines velocity tracking, a joint-torque penalty, a collision penalty, and more, computed at every simulation step.

Related
Sparse Reward · Reward Shaping · Reward Function · Reward Engineering · Reward Hacking · Eureka
Sources
Gymnasium-Robotics: Fetch Reach
Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)

See it in the full glossary →