Visual Matching (SimplerEnv)
视觉匹配VMAdvancedA SimplerEnv evaluation setup that makes the simulated image look as close as possible to real-robot camera footage.
Visual Matching is one of two evaluation setups from SimplerEnv (SIMPLER, released in May 2024 by UC San Diego, Stanford, Berkeley, and Google DeepMind), used to evaluate in simulation manipulation policies that were trained on real-robot data. Such policies often fail as soon as they enter simulation, simply because the images look different. Visual Matching uses “green-screening” to composite the simulated objects and robot arm onto a real background photo, then projects real textures onto the simulated models and adjusts the arm's color to match real footage, bringing the image closer to what a real robot camera sees. The other setup, Variant Aggregation, instead generates multiple variants by changing background, lighting, and distractors and averages across them. Both are checked against real-robot results using the Pearson correlation coefficient and Mean Maximum Rank Violation (MMRV).
ExampleVLA papers evaluating on SimplerEnv's Google-robot tasks — picking up a Coke can, moving near an object, opening and closing a drawer — typically report success rate in two separate columns, “Visual Matching” and “Variant Aggregation.”
- Also called
- VM, SimplerEnv Visual Matching, SIMPLER Visual Matching
- Related
- SimplerEnv · Variant Aggregation (SimplerEnv) · Sim-to-Real Correlation · Mean Maximum Rank Violation · Sim-to-Real Gap (Reality Gap) · Real-to-Sim
- Sources
- Evaluating Real-World Robot Manipulation Policies in Simulation (arXiv 2405.05941)
SIMPLER 项目主页 (Chinese)
SimplerEnv GitHub 仓库 (Chinese) - As of
- 2024-05