Action Expert
动作专家EssentialA separate set of parameters inside a VLA dedicated to producing continuous actions, apart from the vision-and-language backbone.
The term “action expert” became popular through Physical Intelligence's 2024 π0. π0 uses a roughly 3-billion-parameter PaliGemma vision-language model as its backbone, plus a separate, newly initialized Transformer of about 300 million parameters that handles robot state and action tokens; that second set of weights is the action expert. The two share the same attention layers but keep separate weights, similar to a mixture-of-experts with just two experts, and information flows only one way: action tokens can attend to the image and text tokens, but not the reverse, so the backbone doesn't get pulled off its original pretraining. The action expert generates continuous action chunks with flow matching; the paper found that giving state and action tokens their own dedicated weights works better than sharing the backbone's weights. π0.5 and SmolVLA keep this same name, and GR00T N1's action module is a similar design.
ExampleAt inference, π0's PaliGemma backbone first encodes the camera images and language instruction, and the action expert then integrates 10 steps of flow matching to output a continuous action chunk 50 steps long, with a control frequency of up to 50Hz.
- Also called
- Action Expert Module
- Related
- Vision-Language-Action Model · Action Head · Flow Matching · Mixture of Experts · π0 · Backbone Network
- Sources
- π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)
π0 论文 HTML 全文(动作专家参数量与结构) (Chinese)