Action Head
动作头EssentialThe output module attached after a backbone that turns extracted features into concrete robot actions.
The action head follows the same “backbone plus head” split used in vision: the backbone extracts features from images, language, and proprioceptive state, and the head turns those features into whatever output the task needs, which for a robot policy means joint angles, an end-effector pose, or gripper open/close. Three kinds of action heads are common: an MLP head that regresses continuous values directly, trained with mean squared error, which tends to average multiple valid ways of doing something into one; a head that discretizes actions and predicts them as classification over tokens; and a diffusion or flow-matching head, which can express a multimodal action distribution. Octo attaches a lightweight diffusion action head after its Transformer output, and when fine-tuning to a new robot it can simply swap in a new action head to match a different action space; GR00T N1 uses a diffusion Transformer as its action module.
ExampleOcto is pretrained with a diffusion action head predicting a segment of future action; when transferred to a new arm with a different action dimensionality, the pretrained Transformer body is kept, a new action head matching the new action space is attached, and the whole model is fine-tuned together on about 100 demonstrations.
- Also called
- Action Decoder, Policy Head
- Related
- Backbone Network · Diffusion Action Head · Action Expert · Continuous Action Regression · Octo · Action Multimodality
- Sources
- Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)