Embodied AI Glossary中文

Q-Learning

Q 学习Common

A classic model-free reinforcement-learning algorithm that repeatedly corrects its Q-values toward the immediate reward plus the next state's best Q-value.

Q-learning was introduced by Chris Watkins in his 1989 doctoral thesis, with a convergence proof given with Peter Dayan in 1992. It needs no environment model: at every step, it uses “immediate reward plus discount times the next state's maximum Q-value” as a target and nudges the current Q-value toward it, a self-referential update that makes it a form of temporal-difference learning. It's an off-policy algorithm: it explores randomly with ε-greedy while collecting data, but learns the greedy, optimal policy anyway, so old data can be reused repeatedly. Early implementations stored Q-values in a table, which only worked for small discrete problems; DeepMind's DQN replaced the table with a neural network. The max operation tends to bias Q-values upward, and double Q-learning was designed specifically to correct this.

ExampleQT-Opt (Google, 2018) trained a Q-learning-based visual grasping policy on more than 580,000 real grasp attempts, reaching a 96% success rate on objects it had never seen during training.

Also called
Tabular Q-Learning
Related
Q-Function · Temporal-Difference Learning · Deep Q-Network · Off-Policy · Overestimation Bias · QT-Opt
Sources
Wikipedia: Q-learning
Hugging Face Deep RL Course: Introducing Q-Learning
QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

See it in the full glossary →