Embodied AI Glossary中文

EmbodiedBench

Advanced

A benchmark with 4 environments testing how well multimodal large models perform as the “brain” of an embodied agent.

EmbodiedBench was proposed by Rui Yang, Tong Zhang, and colleagues at UIUC and other institutions, published at ICML 2025, specifically to evaluate how well multimodal large language models (MLLMs, models that can both look at images and read text) perform as the “brain” of an embodied agent. It has 4 environments totaling 1,128 test tasks, split into two levels: EB-ALFRED and EB-Habitat test high-level task decomposition and planning (for example, “put the book on the table,” where the model outputs a sequence of high-level skills); EB-Navigation and EB-Manipulation test low-level action planning (directly outputting control values such as translation and rotation), which demands precise perception and spatial reasoning. Tasks are also grouped by 6 capabilities: basic tasks, common-sense reasoning, complex-instruction understanding, spatial awareness, visual perception, and long-horizon planning. The authors tested 24 closed- and open-source models and found they were good at high-level tasks but struggled with low-level manipulation; the paper's abstract reports the best model, GPT-4o, averaging only 28.9%.

ExampleIn EB-Manipulation, a model sees a tabletop image and the instruction “stack the red block on the blue block” and must directly output the arm end effector's position, orientation, and gripper state, rather than calling a ready-made “grasp” skill.

Also called
Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents, EB-ALFRED, EB-Habitat, EB-Navigation, EB-Manipulation
Related
Benchmark · Multimodal Large Language Model · ALFRED · Habitat · LLM-based Task Planning · Embodied Arena
Sources
EmbodiedBench (arXiv 2502.09560)
EmbodiedBench 项目主页 (Chinese)
As of
2025-02

See it in the full glossary →