Embodied AI Glossary中文

Visual Question Answering

视觉问答VQACommon

Given an image and a natural-language question about it, the model answers in words.

Visual question answering takes an image and a question about it, such as “how many cups are on the table?”, and outputs a natural-language answer, requiring the model to understand the image, the language, and relevant common sense together. The task was introduced by Antol, Agrawal, and colleagues at ICCV 2015; the VQA dataset has about 250,000 images and 760,000 questions. The 2017 VQA v2 rebalanced the dataset to reduce cases where a model could guess the answer from the question alone, without looking at the image. Today most vision-language models are trained and evaluated with question-answering formats, and embodied and spatial-reasoning benchmarks like ERQA and VSI-Bench are also framed as question answering; VQA data is often mixed into VLA training as a co-training task — for example, π0.5 uses web data including VQAv2 to preserve its image-understanding ability. Unlike embodied question answering, VQA only looks at a given image and never requires the robot to move around to find the answer.

ExampleGiven a photo of a kitchen counter, ask “is the cup to the left of the sink empty?” and the model answers “yes, it’s empty”; an embodied-reasoning benchmark would instead ask something like “which object should the robot’s gripper move above first?”

Also called
VQA, Image QA
Related
Vision-Language Model · Multimodal Large Language Model · Embodied Question Answering · Co-training · ERQA · VSI-Bench
Sources
VQA: Visual Question Answering (ICCV 2015)
VQA 官网(VQA v2 数据集与挑战赛) (Chinese)
π0.5: a Vision-Language-Action Model with Open-World Generalization(协同训练用到 VQAv2 等网络数据) (Chinese)
As of
2025-04

See it in the full glossary →