Embodied AI Glossary中文

MDM (Motion Diffusion Model)

MDM(人体动作扩散模型)MDMAdvanced

A landmark method that generates a 3D human motion sequence from text or an action category using a diffusion model.

MDM (Human Motion Diffusion Model) is work released in September 2022 by Guy Tevet, Amit Bermano, and colleagues at Tel Aviv University in Israel, published at ICLR 2023. It applies the diffusion models used in image generation to human motion: using a Transformer as the network, it progressively denoises a sequence of joint motion from noise, with classifier-free guidance letting generation be controlled by text or an action category. A key design choice is predicting the clean motion directly at every step rather than the noise, which makes it possible to add geometric losses on joint position, velocity, and foot contact, producing more natural motion. It achieved the best results at the time on text-to-motion benchmarks such as HumanML3D and KIT, and became the baseline for a large amount of follow-up motion-generation work. In embodied AI, the human motion this kind of model generates can, after motion retargeting, be handed to a humanoid robot's motion-tracking controller for execution.

ExampleGiven the prompt “a person walks forward a few steps and then sits down,” MDM generates a corresponding 3D human skeletal motion sequence.

Also called
Human Motion Diffusion Model
Related
Diffusion Model · Human Motion Generation · HumanML3D · Classifier-Free Guidance · Motion Retargeting · Motion Tracking
Sources
arXiv 2209.14916: Human Motion Diffusion Model
MDM 项目主页 (Chinese)
As of
2022-09

See it in the full glossary →