Vision Encoder
视觉编码器EssentialA network that turns an image into a set of feature vectors, or visual tokens, for downstream models to use.
A vision encoder turns pixels into features a model can use: given an image, it outputs a set of vectors, each summarizing the content of a small region. Early designs mostly used convolutional networks such as ResNet; the mainstream today is the Vision Transformer (ViT), which cuts an image into patches — 16×16 pixels in the original ViT — turns each patch into a vector, and uses attention to let the patches exchange information. An encoder's ability mostly comes from pretraining: OpenAI's 2021 CLIP trained on 400 million web image-text pairs to align images and text in the same space; SigLIP is Google's improved version built on that idea; and Meta's DINOv2 trains purely on images with self-supervision, preserving more spatial and geometric detail. VLAs usually use an off-the-shelf vision encoder as is, connecting its features to the language model through a projection layer, and may freeze it or fine-tune it jointly during training.
ExampleOpenVLA feeds a 224×224 image into both SigLIP and DINOv2 at once, concatenates the two sets of features channel-wise, and passes them through a two-layer MLP projection layer to turn them into visual tokens the language model can read.
- Also called
- Image Encoder, Visual Backbone Network
- Related
- Vision Transformer · CLIP · SigLIP · DINOv2 · Projector / Connector · Visual Token
- Sources
- An Image is Worth 16x16 Words (ViT, arXiv:2010.11929)
Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv:2103.00020)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246) - As of
- 2024-06