Embodied AI Glossary中文

LIBERO-PRO

Advanced

Adds object, position, instruction, and environment perturbations to LIBERO to test whether a VLA model is just memorizing.

LIBERO-PRO is an extended evaluation benchmark released in October 2025 by teams from Huazhong University of Science and Technology, Harvard, MIT, Lehigh University, and others. The authors point out that LIBERO's training and test scenes are nearly identical, so a model can score well by memorizing action sequences and tabletop layouts, inflating scores and making fair comparison difficult. LIBERO-PRO adds reasonable perturbations across four dimensions — the manipulated object, initial state, task instruction, and environment — which can be freely combined. Results show that models scoring above 90% success under the standard setup can drop to 0% under the generalization setup: a model keeps reaching for the target object even after it's swapped for an unrelated one, and its actions barely change even when the instruction is scrambled or replaced with gibberish. It appeared around the same time as LIBERO-Plus, and both point to LIBERO having become saturated as a benchmark.

ExampleThe project homepage reports OpenVLA scoring 0.98 success on the original tasks, dropping to 0.00 after object-position perturbation; π0.5 drops from 0.97 to 0.38.

Also called
Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
Related
LIBERO Benchmark · LIBERO-Plus · Benchmark Saturation · Generalization / Robustness Evaluation · Out-of-Distribution · Leaderboard Chasing
Sources
LIBERO-PRO: Towards Robust and Fair Evaluation of VLA Models Beyond Memorization (arXiv 2510.03827)
LIBERO-PRO GitHub 仓库 (Chinese)
As of
2025-10

See it in the full glossary →