Embodied AI Glossary中文

Q-Chunking

动作分块强化学习QCAdvanced

Doing reinforcement learning where both the policy and the Q-function operate on a whole chunk of actions at once.

Q-chunking is a reinforcement learning method proposed by Sergey Levine's team at Berkeley in 2025, aimed at long-horizon, sparse-reward, offline-to-online tasks, where a policy is pretrained on existing data and then goes online. It brings action chunking, a technique common in imitation learning, into reinforcement learning: the policy outputs a chunk of h consecutive actions at once, and the Q-function, which judges how good an action is, also takes the state plus the whole action chunk as input. This has two benefits: acting in chunks makes exploration more coherent and lets the agent carry over behavioral habits from the offline data, and scoring a whole chunk at once enables unbiased multi-step temporal-difference updates, so value information propagates faster. To keep the policy from drifting too far from the data, it samples several action chunks from a flow-matching policy and picks the one with the highest Q-value, or adds a distillation constraint.

ExampleOn OGBench's block- and puzzle-manipulation tasks and on robomimic tasks, with 1 million steps of offline pretraining followed by 1 million steps of online interaction, Q-chunking with a chunk length of 5 clearly beats offline-to-online baselines such as RLPD, and the harder the task, the bigger the gap.

Also called
Reinforcement Learning with Action Chunking, QC
Related
Action Chunking · Offline-to-Online Reinforcement Learning · Temporal-Difference Learning · Reinforcement Learning with Prior Data · Flow Matching · Sparse Reward
Sources
Li, Zhou, Levine 2025: Reinforcement Learning with Action Chunking
As of
2025-07

See it in the full glossary →