Target Network
目标网络AdvancedA slowly updated copy of the main network, used only to compute training targets, making Q-learning more stable.
In temporal-difference methods like Q-learning, the training target is immediate reward plus the discounted Q-value of the next state, and that Q-value is computed by the very same network being trained, effectively chasing a target that moves as you move, which is prone to diverge with neural networks. DeepMind's 2015 DQN paper, published in Nature, introduced the target network: a copy of the main network dedicated to computing targets, synchronized only every fixed number of steps. Continuous-control algorithms such as DDPG, TD3, and SAC switched to soft updates instead, nudging the target network's parameters toward the main network's a little every step, called Polyak averaging, with a coefficient close to 1. Together with experience replay, it is a standard component of off-policy deep reinforcement learning; the TD3 paper also analyzes its relationship to Q-value overestimation.
ExampleIn OpenAI's Spinning Up documentation, DDPG's target network is soft-updated as φ_targ ← ρ·φ_targ + (1−ρ)·φ, with an example ρ of 0.995, while DQN-style algorithms instead copy the whole main network over every fixed number of steps.
- Also called
- Target Q-Network
- Related
- Deep Q-Network · Q-Function · Temporal-Difference Learning · Experience Replay · Overestimation Bias · Soft Actor-Critic
- Sources
- OpenAI Spinning Up: Deep Deterministic Policy Gradient(Target Networks 一节) (Chinese)
Wikipedia: Q-learning(Deep Q-learning 一节) (Chinese)
Fujimoto et al. 2018: Addressing Function Approximation Error in Actor-Critic Methods (TD3)