Embodied AI Glossary中文

Mixture of Experts

混合专家模型MoECommon

A model made of many 'expert' subnetworks where only a few are activated per input, trading a little compute for a lot of capacity.

The idea behind Mixture of Experts goes back to Jacobs, Jordan, Nowlan, and Hinton in 1991: several expert networks plus a gating network (also called a router) that decides which experts handle each input. In 2017, Shazeer and colleagues turned this into a sparsely-gated MoE layer for large-scale language models with hundreds of billions of parameters, where each input only runs through a small fraction of them. It solves a specific problem: you want a bigger model that holds more knowledge, but you don't want every inference to pay for the full parameter count. Large language models such as Mixtral 8x7B, DeepSeek-V3, and Llama 4 all use MoE. Worth noting: the π0 paper describes its own architecture as ‘an MoE with only two experts’ — images and text go through the VLM weights, and state and action go through the action-expert weights — a fixed division of labor by token type, not routing learned by a gating network.

ExampleMixtral 8x7B has 8 experts per layer and routes each token to 2 of them; total parameter count is about 46.7B, but each token actually only uses about 12.9B parameters worth of compute.

Also called
MoE, Sparse MoE
Related
Mixture-of-Transformers · Action Expert · Large Language Model · Parameter Count (Model Size) · Transformer · Multilayer Perceptron
Sources
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (arXiv:1701.06538)
Wikipedia: Mixture of experts
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)

See it in the full glossary →