Embodied AI Glossary中文

RL-based Locomotion Control

强化学习运控Common

Training a neural network with reinforcement learning in simulation to directly control a legged robot's walking.

RL-based locomotion control means using reinforcement learning (maximizing reward through trial and error) to train a neural-network controller that lets quadrupeds, humanoids, and other legged robots walk, run, and get up after falling. A typical pipeline runs thousands of robots in parallel inside a GPU simulator such as Isaac Gym or Isaac Lab, training with PPO; the policy outputs joint target angles at roughly 50 Hz, which per-joint PD controllers convert into torques; domain randomization and actuator modeling then help bridge the sim-to-real gap for deployment on hardware. ETH's Hwangbo et al. validated this approach on the ANYmal quadruped in 2019, and Rudin et al. cut flat-ground walking training down to under 4 minutes on a single GPU in 2021. Compared to model-based methods such as MPC plus whole-body control, it doesn't require hand-written gaits or a precise model and tends to be more robust on complex terrain, but it depends heavily on reward design and extensive tuning.

ExampleIn Unitree's open-source unitree_rl_gym, training Go2 uses a 5-millisecond simulation step with the action updated every 4 steps — so the policy outputs 12 joint target angles every 20 milliseconds — tracked by PD control with Kp=20, Kd=0.5; the trained network is then exported and deployed to the real robot.

Also called
RL Locomotion
Related
Legged Locomotion · Sim-to-Real Transfer · Domain Randomization · Proximal Policy Optimization · Stiffness and Damping Gains · Velocity Command Tracking
Sources
Learning agile and dynamic motor skills for legged robots (Hwangbo et al., Science Robotics 2019)
Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (Rudin et al.)
unitree_rl_gym:Go2 训练配置 (Chinese)
As of
2026-09

See it in the full glossary →