Goal-Conditioned Reinforcement Learning
目标条件强化学习GCRLAdvancedReinforcement learning where both the policy and the value function take the goal as input, letting one model reach many goals.
Ordinary reinforcement learning usually learns just one fixed task; goal-conditioned reinforcement learning writes the policy as π(a|s,g), so the same network carries out different tasks depending on the goal g, with the reward usually just “did it reach the goal or not.” A landmark example is DeepMind's 2015 Universal Value Function Approximators (UVFA), which extends the value function to V(s,g), able to generalize to goals never seen during training. Its biggest difficulty is sparse reward: most attempts never reach the goal, so there's no learning signal, which is why it's often paired with hindsight experience replay (relabeling a failed trajectory as having reached whatever state it actually ended up at). Goals can be coordinates, images, or language, and robot grasping, pushing, and navigation are all commonly framed this way; the low-level policy in hierarchical reinforcement learning is also often goal-conditioned.
ExampleA robot arm pushing a block: the goal is any position on the table, and success means the block ends up near that goal; the same policy learns to push the block to many different positions.
- Also called
- GCRL
- Related
- Goal-conditioned Policy · Hindsight Experience Replay · Sparse Reward · Value Function · Goal-Conditioned Behavior Cloning · Hierarchical Reinforcement Learning
- Sources
- Goal-Conditioned Reinforcement Learning: Problems and Solutions (IJCAI 2022 survey)
Universal Value Function Approximators (ICML 2015)