Automatic Domain Randomization
自动域随机化ADRAdvancedA domain-randomization method, proposed by OpenAI, that automatically widens its randomization range as the policy gets better.
ADR is an algorithm OpenAI proposed in its 2019 work on solving a Rubik's Cube with a robot hand. Ordinary domain randomization (randomly varying parameters like friction, mass, and appearance in simulation so a policy transfers to a real robot) requires a person to manually set the randomization range for every parameter, which becomes hard to tune as the number of parameters grows. ADR turns this into an automatic curriculum: the initial distribution is concentrated entirely on a single environment set to calibrated real-robot values; during training, one parameter is randomly picked and pinned to the current edge of its range for evaluation, and its range is widened if performance is above an upper threshold or narrowed if it is below a lower threshold. Difficulty then grows in step with the policy's ability, and the final range can end up far wider than anything a person would set by hand. The paper used ADR to jointly train a control policy and a vision-based pose-estimation network, both trained entirely in simulation and then transferred to a real Shadow dexterous hand.
ExampleThe paper's best policy achieved about a 60% success rate on real-robot cube configurations requiring 15 moves to solve, and about 20% on the hardest configurations requiring 26 moves; the solution sequence itself came from the classical Kociemba solver, while the neural network handled the hand's manipulation.
- Also called
- ADR
- Related
- Domain Randomization · Dynamics Randomization · Sim-to-Real Transfer · Curriculum Learning · Dactyl · Shadow Dexterous Hand
- Sources
- Solving Rubik's Cube with a Robot Hand (OpenAI, arXiv 1910.07113)
- As of
- 2019-10