Embodied AI Glossary中文

Sparse Reward

稀疏奖励Common

A reward given only at a few key moments, such as task completion, with zero reward the rest of the time.

Sparse reward means the environment gives zero reward almost all the time, with a signal only when the goal is reached — the most typical form is a binary reward: 1 for success, 0 otherwise. The upside is that it's simple to define and hard to game, since it directly reflects whether the task was actually completed; the difficulty is that an agent exploring randomly rarely stumbles into success by chance, so it can go a long time with no learning signal at all, and this gets worse as tasks get longer. Common countermeasures include reward shaping to add intermediate hints (the opposite approach is called dense reward), using human demonstrations to supply examples of success, hindsight experience replay (HER, which relabels a failed trajectory's actual endpoint as the goal, turning it into a “success” example), and curriculum learning from easy to hard. Real-robot reinforcement learning often trains a success detector to generate this kind of 0/1 reward automatically.

ExampleHIL-SERL collects about 200 success images and 1,000 failure images per task through teleoperation and trains a binary classifier as the reward: a positive reward is given only when the classifier judges the task complete, and 0 otherwise; the classifier's accuracy on held-out data is typically above 95%.

Also called
Binary Reward, Success Reward
Related
Dense Reward · Reward Shaping · Hindsight Experience Replay · Success Detector · Exploration vs. Exploitation · HIL-SERL
Sources
Hindsight Experience Replay (arXiv 1707.01495)
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)

See it in the full glossary →