Embodied AI Glossary中文

Querying Transformer

Q-FormerAdvanced

The small module in BLIP-2 that uses 32 learnable queries to distill image information before handing it to a large language model.

Q-Former is a connector module Salesforce introduced in its 2023 BLIP-2 paper, used to bridge a frozen image encoder and a frozen large language model. It's a small Transformer of about 188 million parameters, initialized from BERT-base, that takes in 32 learnable query vectors (768 dimensions each), reads information out of the image features through cross-attention, and outputs a fixed-size 32×768 feature — far smaller than ViT-L/14's raw 257×1024 features, acting as an information bottleneck. Training has two stages: first learning image-text representations alongside the image encoder, then attaching the language model to learn generation. Because only Q-Former and a few other parameters are trained, BLIP-2 beats Flamingo-80B by 8.7% on zero-shot VQAv2 while using 54x fewer trainable parameters. InstructBLIP and others reuse it; LLaVA and Prismatic instead switch to a simpler MLP projector, and OpenVLA's base model uses an MLP too.

ExampleBLIP-2 hands the image features extracted by ViT-g to Q-Former, which compresses them into 32 vectors; after a single linear projection, these serve as a 'soft visual prompt' prepended to the text input of the OPT or Flan-T5 language model.

Also called
Q-Former, BLIP-2 Q-Former
Related
Projector / Connector · Perceiver Resampler · Learnable Query · Cross-Attention · Vision-Language Model · Vision Encoder
Sources
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (arXiv 2301.12597)
BLIP-2 full text (HTML, Section 3.1 Model Architecture)

See it in the full glossary →