Denoising World Model Learning
DWL(去噪世界模型学习)DWLAdvancedAn end-to-end reinforcement-learning walking framework using only proprioception, letting a humanoid cross snow, slopes, and stairs.
DWL was proposed in 2024 by RobotEra together with Jianyu Chen's group at Tsinghua University and the Shanghai Qi Zhi Institute, and was a best-paper finalist at RSS 2024. When a humanoid moves from simulation to a real robot, it faces environmental disturbances, inaccurate dynamics modeling, sensor noise, and quantities like linear velocity or contact force that simply can't be measured directly. DWL treats all of this as noise added to the true state: during simulated training, noise is added to observations and unmeasurable quantities are masked out, a GRU encoder extracts a latent state from a window of history, and a decoder reconstructs the full state — including privileged information like friction coefficient, external force, and terrain height — a kind of “denoising” world model; the policy is trained on the latent state with PPO, while the critic sees the full state directly (asymmetric actor-critic). With no camera or lidar, the policy transfers zero-shot to the 1.2-meter XBot-S and the 1.65-meter XBot-L, walking over snow, slopes, stairs, and rough ground with the same set of parameters.
ExampleIn indoor tests, DWL reached 100% success climbing and descending 10-centimeter-high steps, on slopes, and on uneven ground; PPO with the denoising loss removed managed only 20% success climbing stairs.
- Also called
- DWL, Advancing Humanoid Locomotion: Mastering Challenging Terrains with DWL
- Related
- Asymmetric Actor-Critic · Privileged Information · Sim-to-Real Transfer · Domain Randomization · Blind Locomotion · RobotEra
- Sources
- Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning (arXiv:2408.14472)
- As of
- 2024-08