Prismatic VLMs
Prismatic VLMAdvancedA Stanford/Toyota Research Institute study and open model that systematically compares VLM design choices; OpenVLA's base model.
Prismatic VLM comes from a Stanford University and Toyota Research Institute (TRI) paper by Siddharth Karamcheti and colleagues, published at ICML 2024. The authors built a unified training and evaluation framework (12 benchmarks covering visual question answering, object localization, and challenge sets) to compare vision-language-model design choices one at a time: which vision encoder to use, whether to run a separate alignment-pretraining stage first, and whether the language model should be the base or the instruction-tuned version. The main findings: skipping the separate alignment stage and just doing single-stage training doesn't hurt performance and saves 20-25% of the compute; concatenating DINOv2 and SigLIP features gives a clear boost on localization tasks; and an instruction-tuned language model has no significant advantage. The resulting 7B-13B models beat the contemporary InstructBLIP and LLaVA v1.5, with code and dozens of checkpoints released open-source. Its connection to embodied AI is that OpenVLA is built directly on Prismatic-7B as its base, and the OpenVLA paper credits this fused vision encoder with helping spatial reasoning.
ExampleOpenVLA's base model, Prismatic-7B, is made of a roughly 600-million-parameter fused DINOv2 + SigLIP vision encoder, a 2-layer MLP projector, and the Llama 2 7B language model.
- Also called
- Prismatic-7B
- Related
- OpenVLA · Vision-Language Model · DINOv2 · SigLIP · Projector / Connector · Llama
- Sources
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models (arXiv 2402.07865)
TRI-ML/prismatic-vlms (GitHub)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246) - As of
- 2024-07