RL-100
AdvancedA diffusion-policy-based real-world reinforcement learning framework that reached 1,000-for-1,000 success across eight manipulation tasks.
RL-100 was released in October 2025 by teams from Shanghai Jiao Tong University, Shanghai Qi Zhi Institute, and Tsinghua University (Huazhe Xu and colleagues), published in Science Robotics in 2026. It targets the near-expert-level reliability that home and factory deployment require. The framework is built on a diffusion visuomotor policy and proceeds in three stages — imitation learning from human demonstrations, offline reinforcement learning, and real-world online reinforcement learning — all three sharing a single clipped PPO objective applied over the denoising process, which keeps improvements conservative and stable; consistency distillation then compresses the multi-step denoising into a single step to meet the demands of high-frequency control. Across eight real-robot tasks — pushing objects, bowling, pouring water, folding cloth, screwing in a bolt, juicing, and folding paper boxes, among others — it achieved 1,000 successes out of 1,000 trials, with completion speed matching or exceeding expert teleoperators.
ExampleA juicing robot deployed in a shopping mall served a continuous stream of random customers zero-shot for about seven hours without a single failure.
- Also called
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- Related
- Real-World Reinforcement Learning · Diffusion Policy · Consistency Model · Proximal Policy Optimization · Offline-to-Online Reinforcement Learning · HIL-SERL
- Sources
- RL-100 (arXiv 2510.14830)
RL-100 project page
Tech Xplore: RL-100 framework helps robots refine learned tasks - As of
- 2026-08