Embodied AI Glossary中文

VIMA

Advanced

A Transformer robot agent that uses interleaved text-and-image 'multimodal prompts' to describe manipulation tasks in one unified format.

VIMA was released in October 2022 by researchers at Stanford, NVIDIA, Caltech, and other institutions (Yunfan Jiang, Linxi Fan, Yuke Zhu, Fei-Fei Li, and others), published at ICML 2023. Robot tasks can be specified in many different ways — showing a demonstration to imitate, describing it in language, or giving a goal image. VIMA unifies all of these into a single 'text interleaved with images' multimodal prompt, such as 'put [image of an object] into [image of a container],' with a Transformer reading the prompt and autoregressively outputting actions. The authors also built the VIMA-Bench simulation benchmark: 17 task templates that can procedurally generate thousands of tabletop tasks, more than 600,000 expert trajectories, and a four-level evaluation of increasingly hard generalization. The paper reports up to 2.9x higher success than other designs in the hardest zero-shot setting.

ExampleA VIMA prompt is a sentence with two small images embedded in it: 'put [image of a red block] onto [image of a green plate]'; VIMA finds the matching objects on a simulated tabletop and completes the pick-and-place, and the same model also handles a prompt that first shows a demonstration image and then says 'do it like this.'

Also called
VIMA: General Robot Manipulation with Multimodal Prompts
Related
VIMA-Bench · Language-conditioned Policy · Compositional Generalization · Tabletop Manipulation · Imitation Learning · Transformer
Sources
VIMA (arXiv 2210.03094)
VIMA project page
As of
2023-05

See it in the full glossary →