Embodied AI Glossary中文

Hierarchical Reinforcement Learning

分层强化学习HRLAdvanced

Reinforcement learning split into layers: a high level sets sub-goals, and a low level executes the concrete actions to reach them.

If a long task is learned directly at the level of low-level actions, reward only ever shows up after a very long delay, making credit assignment (figuring out which step actually caused the outcome) extremely hard. Hierarchical reinforcement learning has a high-level policy pick sub-goals or skills at a lower frequency, while a low-level policy outputs concrete actions at a higher frequency to carry them out. Classic frameworks include Sutton and colleagues' 1999 options framework and Dayan and Hinton's feudal reinforcement learning; in the deep-learning era there's DeepMind's 2017 FeUdal Networks (a manager sets goals, a worker outputs actions) and Nachum, Levine, and colleagues' 2018 HIRO (the high level hands the low level a goal, with an off-policy correction for reusing old data). Today's embodied-AI architectures that combine large-model planning with low-level skills, or fast-slow dual systems, follow a similar layered logic, though they aren't always trained with reinforcement learning.

ExampleHIRO, on tasks like a simulated quadruped “ant” robot navigating a maze, has the high level issue a desired state as a sub-goal every fixed number of steps, while the low level controls each leg to reach it.

Also called
HRL
Related
Hierarchical Architecture · Dual-System Architecture (System 1 / System 2) · Skill Primitive · Long-horizon Task · Goal-Conditioned Reinforcement Learning · Credit Assignment
Sources
FeUdal Networks for Hierarchical Reinforcement Learning (arXiv 1703.01161)
Data-Efficient Hierarchical Reinforcement Learning (HIRO, NeurIPS 2018)

See it in the full glossary →