Backbone Network
骨干网络EssentialThe main network that extracts general-purpose features from raw input, with task-specific heads attached after it.
“Backbone network” originally comes from computer vision — 2017's Mask R-CNN paper, for instance, splits the network into a convolutional backbone that extracts features from the whole image, such as a ResNet-50, and heads that do classification, box regression, and mask prediction. A backbone is usually pretrained on large-scale data first and then reused across different tasks, and swapping in a stronger backbone tends to lift performance across the board. In embodied models, the backbone is usually an already-pretrained vision-language model: OpenVLA uses Llama 2 with DINOv2 and SigLIP vision encoders, π0 uses PaliGemma, and GR00T N1 uses NVIDIA's Eagle-2. The backbone supplies semantic knowledge, while an action head or action expert turns its features into actions; whether to freeze the backbone during fine-tuning is a common design choice.
Example1.34 billion of GR00T N1's 2.2 billion parameters belong to the Eagle-2 vision-language backbone; it feeds the action module features from the backbone's 12th layer rather than its last layer, which the paper says makes inference faster and also raises the policy's success rate.
- Also called
- Backbone
- Related
- Vision Encoder · Vision-Language Model · Action Head · Backbone Freezing · Pre-training · Residual Network
- Sources
- Mask R-CNN (arXiv 1703.06870)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)