Embodied AI Glossary中文

MaskedMimic

Advanced

An NVIDIA unified physics-based humanoid controller that treats every control mode as filling in missing motion.

MaskedMimic was released in September 2024 by Chen Tessler, Xue Bin Peng, and colleagues at NVIDIA Research, published at SIGGRAPH Asia 2024 (ACM TOG). Previously, controlling a humanoid character in physics simulation — tracking a motion, walking, reaching for an object, acting out text — usually meant training a separate policy with its own reward for each. MaskedMimic unifies all of these as a motion-inpainting problem: given only partial constraints (target positions for a few joints, some keyframes, a line of text, or an object to interact with), with the rest masked out, the model generates complete, physically plausible whole-body motion. Training happens in two steps: first, reinforcement learning trains a teacher controller on AMASS motion-capture data that fully tracks a reference motion; then DAgger-style behavior cloning distills it into a conditional-VAE student policy that can handle randomly masked input. The code is included in NVIDIA's open-source framework, ProtoMotions.

ExampleGiven only the head and hand target positions corresponding to a VR headset and its two controllers, with every other joint masked out, MaskedMimic generates complete whole-body motion for standing, walking, and reaching.

Also called
Unified Physics-Based Character Control Through Masked Motion Inpainting
Related
Motion Tracking · Perpetual Humanoid Control · DeepMimic · Conditional Variational Autoencoder · DAgger · ProtoMotions (NVIDIA GPU-accelerated humanoid simulation & learning framework)
Sources
arXiv 2409.14393: MaskedMimic
NVIDIA Research: MaskedMimic 项目页 (Chinese)
As of
2024-09

See it in the full glossary →