RL-based Locomotion Control
强化学习运控CommonTraining a neural network with reinforcement learning in simulation to directly control a legged robot's walking.
RL-based locomotion control means using reinforcement learning (maximizing reward through trial and error) to train a neural-network controller that lets quadrupeds, humanoids, and other legged robots walk, run, and get up after falling. A typical pipeline runs thousands of robots in parallel inside a GPU simulator such as Isaac Gym or Isaac Lab, training with PPO; the policy outputs joint target angles at roughly 50 Hz, which per-joint PD controllers convert into torques; domain randomization and actuator modeling then help bridge the sim-to-real gap for deployment on hardware. ETH's Hwangbo et al. validated this approach on the ANYmal quadruped in 2019, and Rudin et al. cut flat-ground walking training down to under 4 minutes on a single GPU in 2021. Compared to model-based methods such as MPC plus whole-body control, it doesn't require hand-written gaits or a precise model and tends to be more robust on complex terrain, but it depends heavily on reward design and extensive tuning.
ExampleIn Unitree's open-source unitree_rl_gym, training Go2 uses a 5-millisecond simulation step with the action updated every 4 steps — so the policy outputs 12 joint target angles every 20 milliseconds — tracked by PD control with Kp=20, Kd=0.5; the trained network is then exported and deployed to the real robot.
- Also called
- RL Locomotion
- Related
- Legged Locomotion · Sim-to-Real Transfer · Domain Randomization · Proximal Policy Optimization · Stiffness and Damping Gains · Velocity Command Tracking
- Sources
- Learning agile and dynamic motor skills for legged robots (Hwangbo et al., Science Robotics 2019)
Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (Rudin et al.)
unitree_rl_gym:Go2 训练配置 (Chinese) - As of
- 2026-09