Meta Reinforcement Learning
元强化学习Meta-RLAdvancedTraining across many similar tasks so an agent can adapt to a new one with only a little trial and error.
Meta reinforcement learning treats “how to do reinforcement learning faster” itself as something to learn: train on a distribution of tasks — say, different payloads or different target positions — so the agent can adapt to a new task from that same distribution using only a handful of interactions. Two approaches are common. Methods like RL² (2016) use a recurrent network that carries memory across episodes, “learning” the new task online through its hidden state. Methods like MAML instead learn a set of initial parameters that can be fine-tuned to a new task in just a few gradient steps. This targets deep RL's poor sample efficiency and weak generalization. In robotics it is commonly used to handle changes in payload, terrain, or body damage, and Meta-World is a common manipulation benchmark built for it.
ExampleNagabandi and colleagues (2018) used meta-learning to train a dynamics model; a real legged mini-robot could use just its last few steps of observation to adapt the model online and keep moving when missing a leg, climbing a slope, or dragging a load.
- Also called
- Meta-RL, Learning to Reinforcement-Learn
- Related
- Meta-Learning · In-Context Learning · Rapid Motor Adaptation · Meta-World · Few-shot · Multi-Task Learning
- Sources
- A Tutorial on Meta-Reinforcement Learning (arXiv:2301.08028)
RL²: Fast Reinforcement Learning via Slow Reinforcement Learning (arXiv:1611.02779)
Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning (arXiv:1803.11347)