Embodied AI Glossary中文

Prismatic VLMs

Prismatic VLMAdvanced

A Stanford/Toyota Research Institute study and open model that systematically compares VLM design choices; OpenVLA's base model.

Prismatic VLM comes from a Stanford University and Toyota Research Institute (TRI) paper by Siddharth Karamcheti and colleagues, published at ICML 2024. The authors built a unified training and evaluation framework (12 benchmarks covering visual question answering, object localization, and challenge sets) to compare vision-language-model design choices one at a time: which vision encoder to use, whether to run a separate alignment-pretraining stage first, and whether the language model should be the base or the instruction-tuned version. The main findings: skipping the separate alignment stage and just doing single-stage training doesn't hurt performance and saves 20-25% of the compute; concatenating DINOv2 and SigLIP features gives a clear boost on localization tasks; and an instruction-tuned language model has no significant advantage. The resulting 7B-13B models beat the contemporary InstructBLIP and LLaVA v1.5, with code and dozens of checkpoints released open-source. Its connection to embodied AI is that OpenVLA is built directly on Prismatic-7B as its base, and the OpenVLA paper credits this fused vision encoder with helping spatial reasoning.

ExampleOpenVLA's base model, Prismatic-7B, is made of a roughly 600-million-parameter fused DINOv2 + SigLIP vision encoder, a 2-layer MLP projector, and the Llama 2 7B language model.

Also called
Prismatic-7B
Related
OpenVLA · Vision-Language Model · DINOv2 · SigLIP · Projector / Connector · Llama
Sources
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models (arXiv 2402.07865)
TRI-ML/prismatic-vlms (GitHub)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)
As of
2024-07

See it in the full glossary →