Generalization / Robustness Evaluation
泛化与鲁棒性评测CommonDeliberately changing lighting, positions, objects, or instructions to test how much skill a policy retains outside its training conditions.
Rather than only reporting a policy's success rate under conditions that match its training distribution, this kind of evaluation systematically applies perturbations — swapping an object's color or shape, adding distractor objects, changing lighting and background, moving the camera viewpoint, changing the robot's starting pose, rewording the language instruction, adding sensor noise — and measures how far the success rate drops. The question it answers is whether a high score reflects real task competence or just memorization of the training scenes. A common approach is to vary one dimension of perturbation at a time to locate specific weaknesses, though multiple perturbations are also stacked together. Notable benchmarks include The Colosseum, LIBERO-Plus, and LIBERO-PRO; the variant aggregations in SimplerEnv follow the same idea. This kind of evaluation routinely shows that models scoring near-perfectly on standard benchmarks lose a large chunk of their success rate under even mild perturbation.
ExampleLIBERO-Plus perturbs LIBERO along seven dimensions — object layout, camera viewpoint, robot initial state, language instructions, lighting, background texture, and sensor noise — and finds that some VLA models' success rates fall from 95% to below 30%, often because the model simply ignores the language instruction.
- Also called
- Perturbation Test, OOD Evaluation, Out-of-Distribution Evaluation, Robustness Benchmark
- Related
- Generalization · Robustness · Out-of-Distribution · Distractor Objects · LIBERO-Plus · The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation
- Sources
- THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation (arXiv 2402.08191)
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models (arXiv 2510.13626) - As of
- 2025-12