Embodied AI Glossary中文

Policy Constraint

策略约束Advanced

In offline RL, keeping the new policy's actions close to what the behavior policy in the dataset actually did.

Policy constraint is a major family of methods in offline reinforcement learning, which trains only on existing data with no further environment interaction. The Q-function, which estimates how much return an action will earn, tends to be overestimated for actions that never appear in the data, and a policy that chases those inflated values falls apart — a failure mode called extrapolation error. Policy constraint methods add a limit: the policy's output actions must stay close to the behavior policy that collected the data, enforced with a distance measure such as KL divergence or MMD, or by adding a behavior-cloning term to the objective. BCQ, BEAR, and TD3+BC all belong to this line; in 2019, Wu, Tucker, and Nachum's BRAC paper unified them under the name “behavior regularization.” Together with conservative-Q-learning-style methods, which instead suppress the value of unfamiliar actions, this forms one of the two main lines of offline RL.

ExampleTD3+BC just adds a behavior-cloning term to the online algorithm TD3's policy update — pulling the output action toward the dataset's actions — and normalizes the states, yet it matched the performance of far more complex offline RL algorithms of its time.

Also called
Behavior Regularization, Behavior Constraint
Related
Offline Reinforcement Learning · Extrapolation Error (OOD Actions in Offline RL) · KL Regularization · Behavior Cloning · Conservative Q-Learning · Advantage-Weighted Regression
Sources
Wu, Tucker, Nachum 2019: Behavior Regularized Offline Reinforcement Learning (BRAC)
Fujimoto & Gu 2021: A Minimalist Approach to Offline Reinforcement Learning (TD3+BC)
Levine et al. 2020: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

See it in the full glossary →