Learnable Query
可学习查询AdvancedA set of vectors trained as model parameters that use attention to gather task-relevant information out of the input.
A learnable query is a set of vectors that are randomly initialized and updated during training, rather than coming from the input; they act as 'queries' in an attention computation that reads the input features, pooling the needed information into a fixed number of outputs. Facebook's 2020 object-detection model DETR uses a set of object queries, each producing one detection; 2023's Q-Former in BLIP-2 uses 32 learnable queries to extract visual features from a frozen image encoder before passing them to a large language model. In robot models, Octo inserts readout tokens into its sequence: they can see the preceding observation and task tokens but aren't seen by those tokens in turn, and the action head generates the action from their output; 2025's VLA-Adapter adds ActionQuery tokens inside a vision-language model (64 of them worked best in its experiments), specifically to gather action-relevant multimodal information for the policy network. Its role is to compress an input of varying length into a fixed-size, task-oriented representation.
ExampleIn BLIP-2, an image is first turned into hundreds of feature tokens by the vision encoder, and Q-Former's 32 query vectors use cross-attention to pool them into 32 outputs, which are projected and prepended to the text before being fed into the language model.
- Also called
- Action Query, Readout Token, Object Query
- Related
- Querying Transformer · Cross-Attention · Perceiver Resampler · Action Head · Octo · VLA-Adapter
- Sources
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and LLMs (arXiv:2301.12597)
Octo: An Open-Source Generalist Robot Policy (arXiv:2405.12213)
VLA-Adapter: An Effective Paradigm for Tiny-Scale VLA Model (arXiv:2509.09372) - As of
- 2025-09