Closed-Loop Evaluation
闭环评测CommonLetting a policy actually control the robot, act, observe, and act again, scored by whether the task finishes.
Closed-loop evaluation means actually running a policy inside a simulated or real environment: at every step, it computes an action from the latest observation, the action changes the environment, and the environment returns a new observation, repeating until the task succeeds, fails, or times out, with success rate and similar metrics tallied at the end. This contrasts with open-loop evaluation, which only compares the model's predicted actions against human demonstrations on offline data (for example, using mean squared error) without the model's own actions ever affecting what it sees next. The problem is that imitation learning's small errors can drive the robot into states never seen in the demonstrations, and these errors compound — something an offline error metric can't reveal. The SimplerEnv paper found in practice that validation-set mean squared error doesn't predict a policy's real-robot performance well, and a CVPR 2024 study in autonomous driving similarly found that open-loop planning metrics can be misleading. VLA papers therefore generally treat closed-loop success rate, in simulation or on a real robot, as the metric that counts.
ExampleEvaluating a VLA on LIBERO: for each task, run several episodes from different initial states with the policy controlling the simulated arm in real time, then report success rate — that's closed-loop evaluation; computing only its predicted actions' mean squared error against a held-out test set would be open-loop evaluation instead.
- Also called
- online evaluation
- Related
- Open-loop Evaluation · Simulation-Based Evaluation · Real-World Evaluation · Rollout · Compounding Error · SimplerEnv
- Sources
- Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving? (arXiv 2312.03031)