Deep Q-Network
深度 Q 网络DQNCommonA reinforcement-learning algorithm that uses a deep neural network to estimate each action's long-term value, its Q-value.
DQN comes from DeepMind's Mnih and colleagues: a 2013 preprint validated it on 7 Atari games, and a 2015 Nature paper extended it to 49 games, using only raw screen pixels and score, reaching performance comparable to a professional human tester overall. It replaces the lookup table in Q-learning (learning how much total return follows from taking a given action in a given state) with a convolutional neural network, and stabilizes training with experience replay (storing past experience and sampling it randomly for training) and a target network (a periodically synced copy used to compute the training target). This launched the deep-reinforcement-learning boom. Because it takes the maximum Q-value over all actions, DQN suits discrete action spaces; for a robot arm's continuous actions, methods like DDPG and SAC are used instead.
ExampleGoogle's QT-Opt (2018) brought Q-learning to real robot-arm grasping: after more than 580,000 real-world grasp attempts, the resulting closed-loop, vision-based grasping policy reached 96% success on objects it had never seen.
- Also called
- DQN, Deep Q-Learning
- Related
- Q-Learning · Q-Function · Experience Replay · Target Network · Deep Deterministic Policy Gradient · QT-Opt
- Sources
- Playing Atari with Deep Reinforcement Learning (arXiv 1312.5602)
Google DeepMind: Deep Reinforcement Learning
QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation (arXiv 1806.10293)