Benchmark Saturation
基准饱和AdvancedWhen leading models' scores on a benchmark all cluster near the maximum, so it can no longer distinguish good methods from bad ones.
A benchmark — a shared test with fixed tasks and an evaluation protocol — tends to saturate the longer it's used: scores from different groups converge toward the ceiling, with gaps shrinking to a percentage point or two, sometimes within the range of random noise. This can happen because methods genuinely improved, but it can equally happen because everyone has repeatedly tuned against the same test set, or because the test scenes are too similar to the training data, letting a model score well through memorization. The Dynabench paper (2021) in NLP made exactly this point: models quickly achieve excellent benchmark scores yet fail on simple adversarial examples. A typical case in embodied AI is LIBERO, where multiple VLA models now score above 90% success under the standard setup. Once a benchmark saturates, the community usually releases a harder or perturbed successor, or shifts to real-robot evaluation and generalization/robustness evaluation. A small lead on a saturated benchmark carries limited weight.
ExampleLIBERO-PRO (2025) shows that a model scoring above 90% success on standard LIBERO drops to 0.0% success once objects are swapped, initial states changed, instructions reworded, or the environment changed — indicating that the high score largely reflected memorization of training trajectories and scene layouts.
- Also called
- Leaderboard Saturation
- Related
- Benchmark · LIBERO Benchmark · LIBERO-PRO · LIBERO-Plus · Leaderboard Chasing · Generalization / Robustness Evaluation
- Sources
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization (arXiv 2510.03827)
Dynabench: Rethinking Benchmarking in NLP (arXiv 2104.14337)
Stanford HAI: The 2025 AI Index Report - As of
- 2025-10