Embodied AI Glossary中文

Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds)

评测统计显著性(置信区间 / 序贯检验 / 多随机种子)Advanced

Using statistics to tell whether a gap in success rate between two policies is real or just noise.

Robot-policy evaluation often runs only a few dozen trials, so the success rate itself carries a wide margin of error: 14 successes out of 20 trials gives roughly a 48%–85% 95% confidence interval by the commonly used Wilson method. Statistically rigorous evaluation means reporting a confidence interval or a Bayesian posterior instead of a single percentage; reinforcement-learning results also need to be repeated across multiple random seeds, since the same algorithm can produce very different outcomes with a different seed. Agarwal and colleagues, at NeurIPS 2021, recommended reporting interval estimates and summarizing multi-task results with interquartile means, and open-sourced the rliable library. Sequential testing runs trials while checking as it goes, stopping early once the gap is already clear enough, saving expensive real-robot trials. The Toyota Research Institute also wrote in 2024 arguing that robot learning should be evaluated to the standards of experimental science: state the experimental conditions clearly, use multiple metrics together, and run proper statistical analysis.

ExampleIn Toyota Research Institute's 2025 large behavior model paper, each real-robot task and policy pair was run 50 times, and each simulated task 200 times, with the evaluator not told which policy was being tested; results were shown as Bayesian posteriors under a Beta prior, pairwise policy comparisons used sequential hypothesis testing, and Bonferroni correction controlled for multiple comparisons.

Also called
Statistically Rigorous Evaluation
Related
Success Rate · Evaluation Protocol · Random Seed and Reproducibility · Double-blind Pairwise Comparison · Real-World Evaluation · Rollout
Sources
Deep Reinforcement Learning at the Edge of the Statistical Precipice (arXiv 2108.13264, NeurIPS 2021)
Robot Learning as an Empirical Science: Best Practices for Policy Evaluation (arXiv 2409.09491)
A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation (TRI, arXiv 2507.05331)
As of
2025-07

See it in the full glossary →