Goal-conditioned Policy
目标条件策略CommonA policy that decides its action based on the goal to reach, not just the current observation.
A goal-conditioned policy also takes a goal as input, usually written π(a | s, g): s is the current state or observation, and g is the goal, which can be a target position, a target image, or a target state. An ordinary policy learns only one fixed task, while a goal-conditioned policy uses a single network to handle a whole family of “reach different goals” tasks, so switching goals requires no retraining. In reinforcement learning this is called goal-conditioned RL, often paired with hindsight experience replay (HER), which relabels a trajectory that failed to reach its original goal as a success toward whatever state it actually reached, easing the sparse-reward problem. In imitation learning, Corey Lynch and colleagues' Play-LMP (2019) trains on unlabeled “play” data with no task labels, and at test time can carry out the corresponding action once given a goal. When the goal is described in a sentence instead, the result is called a language-conditioned policy.
ExampleGiven a robot arm a photo showing “the block in the upper-left corner of the table” as the goal, the policy pushes the block there; given a different photo, the same policy pursues the new goal instead.
- Related
- Goal-Conditioned Reinforcement Learning · Hindsight Experience Replay · Language-conditioned Policy · Goal-Conditioned Behavior Cloning · Hindsight Relabeling · Policy
- Sources
- Goal-Conditioned Reinforcement Learning: Problems and Solutions (IJCAI 2022 Survey, arXiv 2201.08299)
Hindsight Experience Replay (arXiv 1707.01495)
Learning Latent Plans from Play (arXiv 1903.01973)