Best-of-N Sampling
最优 N 采样BoNAdvancedSampling N candidate outputs at once and using a scorer to pick the single highest-scoring one to actually execute.
Best-of-N sampling is a form of test-time compute (spending extra computation at inference time in exchange for better results): a generative model samples N candidates for the same input, and a verifier — a reward model, a value function, or a VLM acting as judge — scores each one, keeping only the highest-scoring candidate. It's widely used with large language models, and its appeal is not needing to retrain the original model at all. Two notable robotics examples: V-GPS (CoRL 2024) reranks a generalist policy's candidate actions using a value function learned with offline reinforcement learning, and the same value function improves 5 different policies across 12 tasks in total; RoboMonkey (CoRL 2025, from Stanford, Berkeley, NVIDIA, and others) samples multiple actions, perturbs them with Gaussian noise and votes, then picks with a trained VLM verifier, improving out-of-distribution task performance by 25 percentage points. The cost is extra forward passes at every step, adding inference latency, and the ceiling on how much it helps depends entirely on how accurate the verifier's scores are.
ExampleAt every control step, RoboMonkey has a VLA model like OpenVLA generate multiple candidate actions for the same frame, has a VLM verifier score each one, and the robot executes only the highest-scoring action.
- Also called
- BoN, Test-Time Verifier, Best-of-N Reranking
- Related
- Inference-Time Compute · Value-Guided Sampling · RoboMonkey · V-GPS · Reward Model · VLM-as-Reward
- Sources
- RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models (arXiv 2506.17811)
RoboMonkey 项目页 (Chinese)
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance (arXiv 2410.13816) - As of
- 2025-09