RoboVQA
RoboVQA 数据集AdvancedA video question-answering dataset from Google DeepMind, built around long-horizon robot tasks.
RoboVQA is a dataset and modeling effort released by Google DeepMind (Pierre Sermanet and others) in November 2023. The dataset contains about 830,000 video-text pairs across 29,500 distinct instructions, with questions centered on long-horizon tasks: what to do next, whether the current step is finished, or whether a certain action can be performed right now. Collection follows a “bottom-up” crowdsourcing approach: open-ended long-horizon tasks are collected first, then executed and annotated by a robot, a person, or a person holding a handheld grasping tool, which the paper reports gives 2.2 times the throughput of traditional step-by-step collection. The authors used it to train a video vision-language model, RoboVQA-VideoCoCa, measuring performance with a unified metric based on the human-intervention rate; the results show that models taking video input have a 19% lower average error rate than models that only see a single image.
ExampleShown a video of a robot in a kitchen and asked “is the task done?” or “what should happen next?”, the model answers in text, for instance “put the sponge in the sink.”
- Also called
- RoboVQA: Multimodal Long-Horizon Reasoning for Robotics
- Related
- Visual Question Answering · Embodied Reasoning · Long-horizon Task · Vision-Language Model · Google DeepMind · Language Annotation
- Sources
- RoboVQA: Multimodal Long-Horizon Reasoning for Robotics (arXiv:2311.00899)
- As of
- 2023-11