Leaderboard Chasing
刷榜CommonOptimizing hard for one benchmark's score in a way that inflates it without necessarily improving real ability.
“Leaderboard chasing” is AI-community slang for researchers or companies repeatedly tuning hyperparameters, cherry-picking settings, or even designing specifically around a public benchmark's quirks to push their score to the top. Chasing a benchmark in moderation can drive real progress, but overdoing it decouples the score from real capability — an instance of Goodhart's Law, which holds that once a measure becomes a target, it stops being a good measure. In embodied AI, a common version of this is simulation-benchmark saturation: several VLA (vision-language-action) models report near-perfect success rates on LIBERO, yet changing the camera angle or adding minor disturbances causes scores to drop sharply. That's why it's worth checking a paper's generalization tests and real-robot evaluations alongside its headline benchmark numbers.
ExampleMany VLA models report above 95% success on LIBERO, while LIBERO-Plus and LIBERO-PRO, which add perturbations, show noticeably lower scores for the same models.
- Also called
- Benchmark Hacking, Gaming the Benchmark
- Related
- Benchmark · Benchmark Saturation · LIBERO Benchmark · LIBERO-Plus · Real-World Evaluation · Cherry-picking
- Sources
- Goodhart's law - Wikipedia
LIBERO Benchmark