Embodied AI Glossary中文

Backbone Network

骨干网络Essential

The main network that extracts general-purpose features from raw input, with task-specific heads attached after it.

“Backbone network” originally comes from computer vision — 2017's Mask R-CNN paper, for instance, splits the network into a convolutional backbone that extracts features from the whole image, such as a ResNet-50, and heads that do classification, box regression, and mask prediction. A backbone is usually pretrained on large-scale data first and then reused across different tasks, and swapping in a stronger backbone tends to lift performance across the board. In embodied models, the backbone is usually an already-pretrained vision-language model: OpenVLA uses Llama 2 with DINOv2 and SigLIP vision encoders, π0 uses PaliGemma, and GR00T N1 uses NVIDIA's Eagle-2. The backbone supplies semantic knowledge, while an action head or action expert turns its features into actions; whether to freeze the backbone during fine-tuning is a common design choice.

Example1.34 billion of GR00T N1's 2.2 billion parameters belong to the Eagle-2 vision-language backbone; it feeds the action module features from the backbone's 12th layer rather than its last layer, which the paper says makes inference faster and also raises the policy's success rate.

Also called
Backbone
Related
Vision Encoder · Vision-Language Model · Action Head · Backbone Freezing · Pre-training · Residual Network
Sources
Mask R-CNN (arXiv 1703.06870)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)

See it in the full glossary →