Embodied AI Glossary中文

Mixture Density Network

混合密度网络MDNAdvanced

A neural network that outputs the parameters of a Gaussian mixture distribution, instead of a single value directly.

The mixture density network was proposed by Christopher Bishop in a 1994 technical report. An ordinary regression network trained with mean squared error learns the average output for a given input; when the same input has several correct answers, that average is often none of them. An MDN instead has the network output the weights, means, and variances of several Gaussian distributions, combined into a Gaussian mixture model, letting it represent a multimodal conditional distribution. One of Bishop's original demonstrations was robot-arm inverse kinematics, where the same end-effector position can correspond to several different sets of joint angles. In robot imitation learning, it's commonly used as a policy's output head to handle action multimodality, as in robomimic's GMM policy; later generative action heads such as Diffusion Policy outperform it on many tasks, but an MDN is cheap to train and run, so it's still often used as a baseline.

ExampleGoing around an obstacle on a table, some demonstrations go left and some go right. A policy trained with mean squared error averages the two into 'drive straight into it'; an MDN outputs two Gaussian components, one for going left and one for right, and just picks one at execution time.

Also called
MDN, GMM Policy Head
Related
Gaussian Mixture Model · Action Multimodality · Gaussian Policy · Continuous Action Regression · Diffusion Policy · Inverse Kinematics (IK)
Sources
Bishop, Mixture Density Networks (Aston University technical report, 1994)
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (robomimic, arXiv:2108.03298)

See it in the full glossary →