Embodied AI Glossary中文

DrEureka

Advanced

Uses a large language model to automatically write reward functions and domain-randomization ranges, transferring a simulation-trained policy to a real robot.

DrEureka is a method proposed in June 2024 by the University of Pennsylvania, NVIDIA, and UT Austin, published at RSS 2024, extending the same team's earlier Eureka (which used a large language model to write reward functions). Moving a policy trained in simulation onto a real robot usually requires a person to repeatedly hand-tune the reward function and the domain-randomization ranges (randomizing physical parameters like friction and mass during training so the policy can tolerate real-world discrepancies). DrEureka automates this in three steps: first, a large language model generates a reward function with safety constraints and trains a policy in simulation with it; then that policy is tested across a range of physical parameters to find the range over which it still works well (reward-aware physics priors, RAPP); finally, the large language model writes a domain-randomization configuration within that range, a final policy is trained on it, and that policy is deployed directly. The simulator used is Isaac Gym, and the policy takes only proprioceptive input.

ExampleOn the Unitree Go1 quadruped, the policy DrEureka trained let the robot dog balance and walk on a yoga ball, with no manual, repeated parameter tuning during training; on standard quadruped forward-walking and dexterous-hand cube-rotation tasks, it performed on par with or better than hand-designed solutions.

Also called
Language Model Guided Sim-To-Real Transfer
Related
Eureka · Sim-to-Real Transfer · Domain Randomization · Reward Function · Reward Engineering · Large Language Model
Sources
DrEureka: Language Model Guided Sim-To-Real Transfer (arXiv 2406.01967)
DrEureka 项目主页 (Chinese)
DrEureka 代码仓库(GitHub) (Chinese)
As of
2024-06

See it in the full glossary →