Embodied AI Glossary中文

Leaderboard Chasing

刷榜Common

Optimizing hard for one benchmark's score in a way that inflates it without necessarily improving real ability.

“Leaderboard chasing” is AI-community slang for researchers or companies repeatedly tuning hyperparameters, cherry-picking settings, or even designing specifically around a public benchmark's quirks to push their score to the top. Chasing a benchmark in moderation can drive real progress, but overdoing it decouples the score from real capability — an instance of Goodhart's Law, which holds that once a measure becomes a target, it stops being a good measure. In embodied AI, a common version of this is simulation-benchmark saturation: several VLA (vision-language-action) models report near-perfect success rates on LIBERO, yet changing the camera angle or adding minor disturbances causes scores to drop sharply. That's why it's worth checking a paper's generalization tests and real-robot evaluations alongside its headline benchmark numbers.

ExampleMany VLA models report above 95% success on LIBERO, while LIBERO-Plus and LIBERO-PRO, which add perturbations, show noticeably lower scores for the same models.

Also called
Benchmark Hacking, Gaming the Benchmark
Related
Benchmark · Benchmark Saturation · LIBERO Benchmark · LIBERO-Plus · Real-World Evaluation · Cherry-picking
Sources
Goodhart's law - Wikipedia
LIBERO Benchmark

See it in the full glossary →