Q-Learning
Q 学习CommonA classic model-free reinforcement-learning algorithm that repeatedly corrects its Q-values toward the immediate reward plus the next state's best Q-value.
Q-learning was introduced by Chris Watkins in his 1989 doctoral thesis, with a convergence proof given with Peter Dayan in 1992. It needs no environment model: at every step, it uses “immediate reward plus discount times the next state's maximum Q-value” as a target and nudges the current Q-value toward it, a self-referential update that makes it a form of temporal-difference learning. It's an off-policy algorithm: it explores randomly with ε-greedy while collecting data, but learns the greedy, optimal policy anyway, so old data can be reused repeatedly. Early implementations stored Q-values in a table, which only worked for small discrete problems; DeepMind's DQN replaced the table with a neural network. The max operation tends to bias Q-values upward, and double Q-learning was designed specifically to correct this.
ExampleQT-Opt (Google, 2018) trained a Q-learning-based visual grasping policy on more than 580,000 real grasp attempts, reaching a 96% success rate on objects it had never seen during training.
- Also called
- Tabular Q-Learning
- Related
- Q-Function · Temporal-Difference Learning · Deep Q-Network · Off-Policy · Overestimation Bias · QT-Opt
- Sources
- Wikipedia: Q-learning
Hugging Face Deep RL Course: Introducing Q-Learning
QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation