Overestimation Bias
Q 值高估AdvancedTaking a max over noisy Q-value estimates systematically inflates them, biasing an agent toward overrated actions.
Q-learning's update target includes a step that takes the maximum Q-value over all actions in the next state. Because Q-values are themselves noisy estimates, taking the max of a set of noisy numbers is biased high in expectation, and that bias then compounds across many rounds of bootstrapping, pushing the agent to favor overrated actions and degrading the policy. Thrun and Schwartz pointed this out as early as 1993. Hado van Hasselt proposed Double Q-learning, using one set of estimates to select the action and another to evaluate it, and in 2015 turned it into Double DQN with DeepMind colleagues, confirming that the original DQN showed clear overestimation on several Atari games. In continuous control, TD3 takes the smaller of two critic networks' values to suppress overestimation, and SAC does the same. In offline RL, overestimation is even worse for actions never seen in the data, which is exactly what methods like conservative Q-learning are built to handle.
ExampleTD3 trains two critic networks at once and, when computing the target value, takes the smaller of the two (clipped double Q-learning), preferring a slight underestimate over an overestimate to make training more stable.
- Also called
- Q-value Overestimation, Maximization Bias
- Related
- Q-Learning · Deep Q-Network · Twin Delayed DDPG · Target Network · Extrapolation Error (OOD Actions in Offline RL) · Conservative Q-Learning
- Sources
- van Hasselt, Guez, Silver 2015: Deep Reinforcement Learning with Double Q-learning
Fujimoto et al. 2018: Addressing Function Approximation Error in Actor-Critic Methods (TD3)