Embodied AI Glossary中文

Update-to-Data Ratio

更新-数据比UTDAdvanced

How many gradient updates are done per step of environment data collected; higher saves data but costs more compute.

The update-to-data ratio, in off-policy reinforcement learning, is how many gradient updates the network does for every one step of interaction added to the experience replay buffer. Standard SAC usually uses 1. Real-robot RL collects data slowly and expensively while compute is comparatively cheap, so there is an incentive to raise this ratio and squeeze more learning out of the same batch of data. But raising it naively tends to make the Q-function overfit and overestimate, hurting training. REDQ (ICLR 2021) used an ensemble of Q-networks to let a model-free algorithm run stably at a UTD far above 1 for the first time; DroQ instead used dropout and layer normalization for a cheaper version. Real-robot RL frameworks such as SERL rely on a high UTD for sample efficiency, and split data collection and training into two separate threads.

Example“A Walk in the Park” (2022) built on SAC for a Unitree A1 quadruped, adding layer normalization and raising the UTD to 20, that is, 20 critic updates per step of data collected, learning to walk from about 20 minutes of real-robot training.

Also called
UTD, UTD Ratio, Replay Ratio
Related
Sample Efficiency · Experience Replay · Off-Policy · Real-World Reinforcement Learning · SERL · Overestimation Bias
Sources
Chen et al. 2021: Randomized Ensembled Double Q-Learning: Learning Fast Without a Model (REDQ, ICLR 2021)
Smith, Kostrikov, Levine 2022: A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning
Luo et al. 2024: SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning

See it in the full glossary →