Embodied AI Glossary中文

LIBERO-Plus

Advanced

A benchmark that adds 7 kinds of perturbations to LIBERO specifically to test the robustness of VLA models.

LIBERO-Plus is a manipulation evaluation benchmark released in October 2025 by teams from Fudan University, the National University of Singapore, Tongji University, and others. It adds 7 categories of controlled perturbation to the popular LIBERO simulation benchmark — object placement, camera viewpoint, robot initial pose, language instructions, lighting, background texture, and sensor noise — for a total of 10,030 task variants, used to test whether a VLA (vision-language-action) model's high score reflects real skill or just memorization of the training scenes. The results show that a small change in viewpoint or initial pose can drop success rate from 95% to below 30%; models turn out to be relatively insensitive to language perturbation, often because they don't really look at the instruction at all. The repository is a drop-in replacement for the original LIBERO, and perturbed training data has also been open-sourced.

ExampleOpenVLA scores 76.5% success on the original LIBERO but falls to 1.1% once the camera viewpoint is changed; replacing OpenVLA-OFT's instructions with blank text barely moves its success rate on the object task group.

Also called
In-depth Robustness Analysis of Vision-Language-Action Models
Related
LIBERO Benchmark · LIBERO-PRO · Generalization / Robustness Evaluation · Robustness · Vision-Language-Action Model · OpenVLA-OFT
Sources
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models (arXiv 2510.13626)
LIBERO-plus GitHub 仓库 (Chinese)
As of
2025-10

See it in the full glossary →