Embodied AI Glossary中文

Reset-Free Reinforcement Learning

无重置强化学习Advanced

Letting a robot keep learning through continuous interaction, without a person resetting the environment every episode.

Standard reinforcement learning assumes the environment resets to its initial state after every episode, trivial in simulation, but on a real robot it usually means a person has to put the objects back, which is one of the main bottlenecks of real-robot RL. Reset-free reinforcement learning studies how to learn with as little manual resetting as possible. Eysenbach and colleagues' 2017 “Leave no Trace” learns a forward policy and a reset policy together, using the reset policy's value function to tell when the agent is about to enter an unrecoverable state. Gupta and colleagues (2021) had multiple tasks reset each other, treating it as a multi-task learning problem. That same year, Sharma, Finn, and colleagues formalized this as “autonomous reinforcement learning” and released the EARL benchmark, finding that ordinary episodic algorithms perform noticeably worse once manual intervention is reduced. It is closely related to failure recovery and autonomous data collection.

ExampleTraining an arm to open a drawer while also learning to close it: after opening, the “close” policy restores the environment, so the robot can practice continuously without anyone standing by to reset it.

Also called
Autonomous RL, Autonomous Reinforcement Learning
Related
Real-World Reinforcement Learning · Episode · Failure Recovery · Autonomous Data Collection · Multi-Task Learning · Online Reinforcement Learning
Sources
Eysenbach et al. 2017: Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
Gupta et al. 2021: Reset-Free Reinforcement Learning via Multi-Task Learning
Sharma et al. 2021: Autonomous Reinforcement Learning: Formalism and Benchmarking

See it in the full glossary →