Embodied AI Glossary中文

GMT

Advanced

A whole-body control method that uses one unified policy to make a humanoid track many different human motions.

GMT was proposed in June 2025 by Xiaolong Wang's group at UC San Diego together with Xue Bin Peng at Simon Fraser University. Motion tracking means having a robot imitate a reference motion — such as human motion-capture data — in real time. Earlier methods often needed one policy per motion category, or per-category fine-tuning; GMT instead uses a single policy that covers walking, kicking, soccer kicks, dancing, and more. Two ideas are key: adaptive sampling, which automatically practices harder clips more during training, and a motion mixture-of-experts (MoE), which lets different parts of the network specialize in different types of motion. Training first uses PPO to produce a teacher policy with access to privileged information, then distills it with DAgger into a student policy that sees only proprioception. The data comes from about 33 hours of motion drawn from AMASS and LAFAN1, deployed on a real 23-degree-of-freedom Unitree G1.

ExampleThe same GMT policy on a Unitree G1 can perform martial-arts kicks and soccer kicks, and also imitate a drunken walk, a crouching walk, and dancing, all without training separately for each motion category.

Also called
General Motion Tracking
Related
Motion Tracking · Mixture of Experts · Unitree G1 · Teacher-Student Distillation · AMASS (Archive of Motion Capture as Surface Shapes) · BeyondMimic
Sources
GMT (arXiv:2506.14770)
GMT 项目主页 (Chinese)
As of
2025-09

See it in the full glossary →