Embodied AI Glossary中文

Conservative Q-Learning

保守 Q 学习CQLAdvanced

Deliberately pushing down the Q-values of actions absent from the dataset, so offline reinforcement learning isn't misled by inflated estimates.

CQL was introduced by Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine in 2020 (NeurIPS 2020), and is a landmark algorithm for offline reinforcement learning (using only a fixed dataset, with no further environment interaction). Ordinary Q-learning fails when applied offline as-is: the policy tends to favor actions absent from the data, whose Q-values (the estimated long-term return of taking a given action in a given state) were never corrected by real data and are often overestimated, so the policy drifts toward these inflated values. CQL adds a regularization term on top of the standard Bellman error: it pushes down the Q-values of actions the policy is likely to pick, while pushing up the Q-values of actions actually taken in the dataset, so the learned value becomes a lower bound on the true value. It's easy to bolt onto existing deep Q-learning or actor-critic algorithms, and the paper reports final returns often 2–5 times those of prior offline methods. Later methods like Cal-QL build directly on it.

ExampleGoogle DeepMind's Q-Transformer (2023) represents a multi-task robot Q-function with a Transformer, training on offline real-robot data combining human demonstrations and autonomously collected data, using an adapted version of CQL's conservative regularizer.

Also called
CQL
Related
Offline Reinforcement Learning · Calibrated Q-Learning · Q-Function · Overestimation Bias · Extrapolation Error (OOD Actions in Offline RL) · Q-Transformer
Sources
Conservative Q-Learning for Offline Reinforcement Learning (arXiv 2006.04779)
NeurIPS 2020 Proceedings: Conservative Q-Learning for Offline Reinforcement Learning
Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions (arXiv 2309.10150)

See it in the full glossary →