OpenEQA (Open-Vocabulary Embodied Question Answering Benchmark)
OpenEQA 开放词汇具身问答基准AdvancedA Meta benchmark testing whether an agent can answer natural-language questions using what it has observed of a real environment.
OpenEQA was released by Meta FAIR in April 2024, with the paper appearing at CVPR 2024. Embodied Question Answering (EQA) requires an agent to understand an environment it currently occupies or has previously seen, and answer questions about it in natural language — for example, “where did I leave my keys?” It contains more than 1,600 question-answer pairs, hand-written rather than generated from templates, drawn from videos and scans of over 180 real environments sourced from HM3D and ScanNet. It defines two settings: episodic memory (EM-EQA), where the agent answers from a previously recorded history of observations, modeling something like smart glasses; and active exploration (A-EQA), where the robot has to move around itself to gather the information it needs. Because answers are open-vocabulary free text, the authors score them automatically with an LLM-based metric called LLM-Match, which they show correlates highly with human judgment.
ExampleIn results Meta published, GPT-4V scored 48.5% accuracy versus 85.9% for humans; on questions that required spatial understanding, models that could see images performed barely better than models that could only read text.
- Also called
- OpenEQA, EM-EQA, A-EQA
- Related
- Embodied Question Answering · Visual Question Answering · Open-vocabulary · Embodied Memory · Habitat-Matterport 3D Dataset · ScanNet
- Sources
- Meta AI Blog: OpenEQA: From word models to world models
OpenEQA project page
GitHub: facebookresearch/open-eqa (data) - As of
- 2024-04