Embodied AI Glossary中文

V-GPS

Advanced

A method that re-ranks a generalist policy's candidate actions at deployment time using a value function learned with offline reinforcement learning.

V-GPS was released by Mitsuhiko Nakamoto, Sergey Levine, and colleagues at UC Berkeley and CMU in October 2024, published at CoRL 2024. Generalist robot policies are trained on demonstration data of wildly varying quality, and the larger the dataset, the harder it is to clean up. V-GPS leaves the original policy untouched: it first trains a Q-function (which scores a state-action pair) with offline reinforcement learning (using only existing data, mainly with Cal-QL) on BridgeData V2 and RT-1 data; at deployment, the generalist policy samples K candidate actions at once (the paper tries 10 and 50), and the Q-function picks the highest-scoring one to execute. It requires no fine-tuning and no access to the policy's weights, and the same value function improves five different policies — Octo, RT-1-X, OpenVLA, and others — making it an early example of trading extra inference-time compute for better performance in robotics.

ExampleOn a real robot, having Octo pick up sushi and put it in a bowl: Octo first generates several candidate actions at each step, V-GPS's Q-function scores each one, and only the highest-scoring action is executed; the project page reports a clear success-rate increase on this kind of task.

Also called
Value-Guided Policy Steering, Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
Related
Value-Guided Sampling · Calibrated Q-Learning · Offline Reinforcement Learning · Q-Function · Inference-Time Compute · Octo
Sources
Steering Your Generalists (arXiv 2410.13816)
V-GPS project page
As of
2024-10

See it in the full glossary →