Embodied AI Glossary中文

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

VLABenchAdvanced

A Fudan University language-conditioned manipulation benchmark focused on common sense, implicit intent, and long-horizon multi-step reasoning.

VLABench is an open-source benchmark released by Fudan University's OpenMOSS team in December 2024 and accepted by ICCV in 2025, evaluating the ability to manipulate a robot arm from natural-language instructions. Built on MuJoCo and dm_control, it uses a 7-DoF Franka arm by default, across 100 task categories (60 atomic, 40 compositional) with more than 2,000 object assets. Compared with earlier benchmarks whose instructions are mostly fixed templates, its tasks require common sense and world knowledge, instructions carry implicit intent, and long-horizon tasks require multi-step reasoning; both VLA policies and VLM-driven workflows can be evaluated on it. The project also provides automatically generated training data and six evaluation tracks.

ExampleIts instructions don't always state directly what to pick up — they carry implicit intent, so the model first has to infer the target object before acting; the paper's results show that even the strongest pretrained VLAs and VLM-based pipelines at the time struggled with these tasks.

Related
Vision-Language-Action Model · Instruction Following · Long-horizon Task · LIBERO Benchmark · MuJoCo (Multi-Joint dynamics with Contact) · Benchmark
Sources
VLABench (arXiv 2412.18194)
VLABench GitHub 仓库(OpenMOSS) (Chinese)
As of
2025-06

See it in the full glossary →