Mean Maximum Rank Violation
平均最大排名违背MMRVAdvancedA metric for whether a simulated evaluation's ranking of policies matches the real-robot ranking; lower is better.
MMRV was proposed in the 2024 SIMPLER (SimplerEnv) paper. The authors argue that a simulated evaluation doesn't need to reproduce the absolute value of real-robot success rate — what matters is getting the relative ranking of policies right. It's computed as follows: for any pair of policies, if their ordering in simulation is reversed compared to the real robot, that counts as one “rank violation,” with a magnitude equal to the difference between their real-robot success rates; each policy takes its single worst violation, and these are averaged across all policies, giving a value between 0 and 1. This way, two policies that were already close in real-robot performance getting swapped only counts as a small error, while swapping two policies with a large real-robot gap counts as a large error. It's usually reported alongside the Pearson correlation coefficient, which only captures linear relationships and can be thrown off by real-robot evaluation noise when policies are close in skill.
ExampleThe SIMPLER paper ranks 6 Google Robot policies: ranking by validation-set action MSE gives an average MMRV of 0.375, while ranking with SIMPLER's “visual matching” simulated evaluation brings it down to 0.056.
- Also called
- MMRV
- Related
- SimplerEnv · Sim-to-Real Correlation · Visual Matching (SimplerEnv) · Variant Aggregation (SimplerEnv) · Simulation-Based Evaluation · Real-World Evaluation
- Sources
- Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)
- As of
- 2024-05