AutoEval
AdvancedBerkeley's round-the-clock, unattended real-robot evaluation system, which judges success automatically and resets the scene itself.
AutoEval was proposed by Zhiyuan Zhou, Pranav Atreya, and colleagues in Sergey Levine's group at Berkeley (one author is also affiliated with NVIDIA), posted to arXiv in March 2025, with its code repository noting publication at CoRL 2025. It automates the two most labor-intensive parts of real-robot evaluation: a fine-tuned PaliGemma vision-language model serves as a success detector, answering questions like “is the drawer open?”; and OpenVLA fine-tuned on a few dozen teleoperated demonstrations (or a recorded trajectory played back) serves as a reset policy that restores the scene to its original state. Users submit their own policy server through a webpage, much like submitting a job to a compute cluster, and the system queues it up to run on a WidowX arm in Bridge-style scenes. The paper reports an average Pearson correlation of 0.942 with human evaluation, about 850 episodes run per station in 24 hours with only 3 human interventions needed, a reduction in human time of over 99%, and it has opened public scenes to the community.
ExampleA researcher deploys their VLA policy as a publicly reachable server, submits an evaluation for the task “put the eggplant in the basket” on the AutoEval webpage, and the system automatically runs several dozen episodes before returning a success-rate report.
- Also called
- Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
- Related
- Real-World Evaluation · Success Detector · BridgeData V2 · OpenVLA · RoboArena · Sim-to-Real Correlation
- Sources
- AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World (arXiv 2503.24278)
AutoEval project page
GitHub: zhouzypaul/auto_eval - As of
- 2025-09