Embodied AI Glossary中文

Value-Guided Sampling

价值引导采样Advanced

Sampling several candidate actions from a policy, then using a value function to score and pick the best one.

Value-guided sampling improves a policy at inference time without touching its weights: at every step, sample several candidate actions from the policy, then score them with a separately trained value function, the Q-function, either taking the highest-scoring one or sampling via softmax over the scores. The representative work is Sergey Levine's group's V-GPS (CoRL 2024): it trains a language-conditioned Q-function with offline RL methods such as Cal-QL on the Bridge and RT-1 datasets, and uses it to rerank five different generalist policies, including Octo and OpenVLA, improving results across 12 tasks. Generalist policies are trained on data of mixed quality, so a value function's job is to pick out the actions that are actually good. This belongs to the same family of inference-time-compute techniques as best-of-N sampling and RoboMonkey's VLM-based action verification.

ExampleIn V-GPS's real-robot experiments, 50 candidate actions are sampled from the generalist policy at every step, and the Q-function picks the highest-scoring one to execute; the paper reports a relative improvement of 82.8% in average success rate across 6 tasks on a WidowX arm.

Also called
Value-Guided Policy Reranking, Value-Guided Policy Steering
Related
Inference-Time Compute · Best-of-N Sampling · V-GPS · Q-Function · Offline Reinforcement Learning · RoboMonkey
Sources
Nakamoto et al. 2024: Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance (V-GPS, CoRL 2024)
Kwok et al. 2025: RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
As of
2025-07

See it in the full glossary →