Embodied AI Glossary中文

Return Conditioning

回报条件化Advanced

Feeding a policy the return you want it to achieve, so it acts toward that target return.

Return conditioning rewrites reinforcement learning as supervised learning: during training, the remaining cumulative reward from each step onward in a trajectory, the return-to-go, is fed into the policy alongside the observation, and the policy is trained to imitate the action actually taken; at test time, a high target return is given, hoping the policy produces high-return behavior. Schmidhuber's 2019 Upside-Down RL and 2021's Decision Transformer are representative: the latter arranges return, state, and action into a sequence for a Transformer to predict actions from, matching mainstream methods on offline RL benchmarks. It needs no learned value function, so training is stable, but Brandfonbrener and colleagues (2022) showed it can fail to find the optimal policy when the environment is very stochastic or data coverage is poor. The advantage conditioning used in π*0.6's RECAP is a variant of the same idea.

ExampleAt test time, Decision Transformer is given a target return, such as 1 for task success, and after every step it subtracts the actual reward received from the target, then predicts the next action conditioned on the remaining target return.

Also called
Return-Conditioned Policy, Return-Conditioned Supervised Learning (RCSL)
Related
Decision Transformer · Advantage Conditioning · Return · Offline Reinforcement Learning · RECAP · Goal-conditioned Policy
Sources
Chen et al. 2021: Decision Transformer: Reinforcement Learning via Sequence Modeling
Schmidhuber 2019: Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Brandfonbrener et al. 2022: When does return-conditioned supervised learning work for offline reinforcement learning?

See it in the full glossary →