Vision-Language Model
视觉语言模型VLMEssentialA model that takes in both images and text and answers in text; the base that most VLAs build on.
A vision-language model takes in an image and text and outputs text: it can describe a picture, answer questions about it, or point out where an object is. Three parts are common: a vision encoder turns the image into features, a projection layer aligns those features to the language model's input space, and the language model handles understanding and generating the text. 2023's LLaVA, for instance, is a CLIP encoder plus a projection layer plus the Vicuna language model; Google's 2024 open-source PaliGemma combines a SigLIP encoder with Gemma-2B, at about 3 billion parameters. The objects, common sense, and spatial knowledge a VLM picks up from internet image-text data are exactly what robots lack, which is why most VLAs start from a VLM: RT-2 adds robot data on top of PaLI-X and PaLM-E through joint fine-tuning, and π0 is built on PaliGemma. VLMs are also commonly used on their own as the high-level planner in a hierarchical architecture.
ExampleAsk a VLM about a kitchen photo, “what's to the left of the sink,” and it answers in text; swap the output for action tokens and train it further on robot data, and you get a VLA like RT-2.
- Also called
- VLM
- Related
- Vision-Language-Action Model · Vision Encoder · Projector / Connector · Large Language Model · PaliGemma · Multimodal Large Language Model
- Sources
- Vision Language Models Explained (Hugging Face blog)
PaliGemma: A versatile 3B VLM for transfer (arXiv:2407.07726)
RT-2: Vision-Language-Action Models (project page) - As of
- 2024-07