Embodied AI Glossary中文

Progress Reward Model

进度奖励模型Advanced

A model that looks at the current frame and estimates how much of a task is done, used as a reward.

A progress reward model is a class of robot reward model: given a task instruction and the current frame, sometimes also the start frame or recent history, it outputs how far along the task is, usually normalized to between 0 and 1. Real manipulation tasks often only offer a sparse succeeded-or-not reward, which reinforcement learning struggles to learn from; a progress score can serve as a dense reward instead, letting the policy know at every step whether it is getting closer to or farther from the goal, and it can also be used to filter data or judge success. There are two main approaches: having a vision-language model estimate progress zero-shot, as in Google DeepMind's GVL, or training one specifically on large amounts of trajectory data, as in 2026's Robometer. The risk is reward hacking, where the policy learns to make the footage merely look like it is progressing.

ExampleGVL shuffles the frame order of a robot video and asks a VLM to estimate the completion percentage of each frame; with no training at all, it gives usable progress values across more than 300 real tasks, and is used for data filtering and success detection.

Also called
Progress Estimator, Task Progress Prediction
Related
Reward Model · Sparse Reward · Dense Reward · Success Detector · Generative Value Learning (GVL) · Robometer
Sources
Ma et al. 2024: Vision Language Models are In-Context Value Learners (GVL)
Liang et al. 2026: Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
Ayalew et al. 2024: PROGRESSOR
As of
2026-03

See it in the full glossary →