Embodied AI Glossary中文

Open-loop Evaluation

开环评测Common

Comparing a model's predicted actions against recorded demonstration actions on offline data, without letting the policy actually control anything.

Open-loop evaluation does not let a policy actually drive a robot. Instead, recorded observations from an offline dataset are fed in frame by frame, and the model's predicted actions are compared against the demonstrated actions, typically using mean squared error (MSE) or L2 error. It is cheap, fast, reproducible, and needs neither a robot nor a simulator. The problem is that the policy's output never affects the next frame, so it cannot capture error accumulation or how well the policy recovers from a mistake; and because the same situation often has more than one correct way to act (action multimodality), a valid but different action still gets penalized simply for not matching the one recorded demonstration. The SIMPLER paper (2024) compared 6 checkpoints and found the Pearson correlation between validation-set MSE and real-robot success rate was only 0.308, versus 0.924 for closed-loop simulated evaluation. Self-driving research has reached a similar conclusion: a CVPR 2024 paper found that on nuScenes' open-loop planning metric, using nothing but the ego vehicle's own state, such as its speed, already scores competitively. As a result, open-loop metrics are mostly used for early debugging and screening.

ExampleOn a validation set from DROID or self-collected data, each frame's image and instruction is fed into a VLA model, and the mean squared error between its output action and the demonstrated action is computed as a rough check for whether training has gone off track.

Also called
Offline Metric Evaluation, Offline Evaluation
Related
Closed-Loop Evaluation · Open-loop Control · Simulation-Based Evaluation · Real-World Evaluation · SimplerEnv · Sim-to-Real Correlation
Sources
Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving? (arXiv 2312.03031)

See it in the full glossary →