Embodied AI Glossary中文

Projector / Connector

投影层Common

A small network that maps features from a vision encoder into the input space of a language model.

A projector is a small network in a multimodal model that connects two separately pretrained components — most commonly, it maps the image features a vision encoder produces into vectors a large language model can read directly. It's needed because the two sides were pretrained independently, so their feature dimensions and distributions don't match. The original LLaVA used just a single linear layer; LLaVA-1.5 switched to a two-layer MLP (multilayer perceptron) and got better results. More elaborate designs, such as a Q-Former or a Perceiver Resampler, also compress the number of tokens along the way. When training a VLM, it's common to freeze both pretrained sides and train only the projector, which is the modality-alignment stage. The same idea shows up in VLAs for robot state and action too: a linear layer or MLP projects low-dimensional vectors like joint angles into the model's embedding dimension.

ExampleOpenVLA concatenates SigLIP and DINOv2 image features channel-wise, then passes them through a two-layer MLP projector before feeding them into the Llama 2 7B language model.

Also called
Projector, Connector, Vision-Language Connector
Related
Vision Encoder · Multimodal Fusion · Querying Transformer · Perceiver Resampler · Modality Alignment (Alignment Pretraining Stage) · LLaVA
Sources
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)
Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, arXiv 2310.03744)

See it in the full glossary →