Benchmark
基准测试EssentialA fixed set of tasks, data, and scoring rules that let different methods be compared under the same conditions.
The term “benchmark” originates in computing, referring to a standardized set of tests used to measure the relative performance of something. In embodied AI, a benchmark typically bundles a fixed set of tasks, a simulator or real-world setup, demonstration data (if any), an evaluation protocol, and a metric — usually success rate. Common simulated benchmarks include LIBERO, CALVIN, SimplerEnv, and RoboTwin, alongside real-robot evaluation networks such as RoboArena. Its value is reproducibility and the ability to compare methods head to head; the risk is that methods can be tuned specifically to score well on the leaderboard (“benchmark hacking”) without that reflecting real-world usefulness, and once leading methods get close to a perfect score, the benchmark stops being able to distinguish between them — what's called benchmark saturation.
ExampleVLA papers commonly report average success rate across LIBERO's four task suites, placing their numbers in the same table as baselines such as OpenVLA for comparison.
- Also called
- leaderboard
- Related
- Baseline · Evaluation Protocol · Success Rate · LIBERO Benchmark · Benchmark Saturation · Leaderboard Chasing
- Sources
- Wikipedia: Benchmark (computing)
LIBERO (NeurIPS 2023 Datasets and Benchmarks Track)