Embodied AI Glossary中文

Simulation-Based Evaluation

仿真评测Essential

Running a policy through many tasks in a simulator and tallying success rate, instead of or alongside real-robot testing.

Simulation-based evaluation means placing a trained robot policy in a simulator and running it through tasks under fixed initial conditions and judging rules, tallying success rate and other metrics. Real-robot evaluation requires a person to place objects and reset the scene by hand, so running dozens or hundreds of trials per policy is slow, and setups differ enough between labs that results are hard to reproduce; simulated evaluation, by contrast, can run automatically, at scale, and repeatably, and is commonly used to compare algorithms, run ablation studies, and pick checkpoints. Common benchmarks include LIBERO, CALVIN, SimplerEnv, and RoboTwin. The main concern is that a simulated score doesn't always reflect real-robot performance — visuals, physics, and the controller all differ from reality, and quite a few benchmarks are already close to saturated — so researchers check the correlation between simulated and real-robot scores, and use real-robot evaluation or world-model-based evaluation as a supplement.

ExampleThe OpenVLA-OFT paper runs simulation-based evaluation on LIBERO's four task suites, raising OpenVLA's average success rate from 76.5% to 97.1%.

Also called
Sim Evaluation
Related
Real-World Evaluation · Benchmark · Success Rate · SimplerEnv · LIBERO Benchmark · Sim-to-Real Correlation
Sources
Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER)
LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT)
As of
2025-02

See it in the full glossary →