Embodied AI Glossary中文

Vision Foundation Model

视觉基础模型VFMCommon

A general-purpose model pretrained on huge amounts of images, usable for many vision tasks with little or no adaptation.

Foundation model refers to a model trained on large-scale data that can adapt to a wide range of downstream tasks, a term a Stanford team coined in 2021; a vision foundation model is the image-focused kind. Three approaches are common. OpenAI's CLIP uses contrastive learning on 400 million image-text pairs from the web to learn visual features aligned with language; Meta's DINOv2 uses self-supervised training on 142 million curated images, producing features especially good at spatial and geometric detail; and Meta's SAM (Segment Anything Model) can segment any object given a point or box prompt. Most of these are built on ViT. Embodied AI rarely trains its vision component from scratch — instead, it typically attaches an off-the-shelf VFM as the vision encoder and feeds its features into a language model or policy network.

ExampleOpenVLA feeds the same camera image into two vision foundation models, SigLIP and DINOv2, concatenates the two sets of features, and projects them into Llama 2's input space with a two-layer MLP.

Also called
VFM, Visual Foundation Model
Related
Foundation Model · Vision Encoder · Vision Transformer · CLIP · DINOv2 · Segment Anything Model
Sources
On the Opportunities and Risks of Foundation Models (arXiv 2108.07258)
DINOv2: Learning Robust Visual Features without Supervision (arXiv 2304.07193)
Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv 2103.00020)

See it in the full glossary →