Reference State Initialization
参考状态初始化RSIAdvancedIn motion-imitation training, starting each episode from a random point in the reference motion instead of always frame one.
Reference state initialization is a reinforcement-learning training trick introduced by Peng and colleagues in 2018's DeepMimic, for training a simulated character or robot to imitate a reference motion such as motion-capture data. If every episode starts from the beginning of the motion, the policy can only learn the first half before the second; for a backflip, without first learning how to land, the early takeoff phase alone tends to make the agent fall and score worse, so it never even gets to practice the landing. RSI instead picks a random point in the reference motion for each episode and sets the character's root position, orientation, velocity, and joint state directly to that frame, letting every phase of the motion get practiced in parallel. It is commonly paired with early termination, ending the episode on a fall, and has become standard practice for training humanoid motion tracking; BeyondMimic goes further and adaptively picks the starting point based on failure rate.
ExampleIn DeepMimic's ablation study, a backflip policy trained without RSI never learns the full flip and only manages a small backward hop; ASAP also used RSI when training a Unitree G1 to imitate highly dynamic motions such as Cristiano Ronaldo's signature jump celebration.
- Also called
- RSI
- Related
- DeepMimic · Early Termination · Motion Tracking · ASAP · BeyondMimic · Exploration vs. Exploitation
- Sources
- Peng et al. 2018: DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
He et al. 2025: ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
Liao et al. 2025: BeyondMimic