Progress Reward Model
进度奖励模型AdvancedA model that looks at the current frame and estimates how much of a task is done, used as a reward.
A progress reward model is a class of robot reward model: given a task instruction and the current frame, sometimes also the start frame or recent history, it outputs how far along the task is, usually normalized to between 0 and 1. Real manipulation tasks often only offer a sparse succeeded-or-not reward, which reinforcement learning struggles to learn from; a progress score can serve as a dense reward instead, letting the policy know at every step whether it is getting closer to or farther from the goal, and it can also be used to filter data or judge success. There are two main approaches: having a vision-language model estimate progress zero-shot, as in Google DeepMind's GVL, or training one specifically on large amounts of trajectory data, as in 2026's Robometer. The risk is reward hacking, where the policy learns to make the footage merely look like it is progressing.
ExampleGVL shuffles the frame order of a robot video and asks a VLM to estimate the completion percentage of each frame; with no training at all, it gives usable progress values across more than 300 real tasks, and is used for data filtering and success detection.
- Also called
- Progress Estimator, Task Progress Prediction
- Related
- Reward Model · Sparse Reward · Dense Reward · Success Detector · Generative Value Learning (GVL) · Robometer
- Sources
- Ma et al. 2024: Vision Language Models are In-Context Value Learners (GVL)
Liang et al. 2026: Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
Ayalew et al. 2024: PROGRESSOR - As of
- 2026-03