Embodied AI Glossary中文

RoboMonkey

Advanced

A method that samples several candidate action sets at deployment time and uses a VLM-based verifier to pick the best one, improving VLA robustness.

RoboMonkey was released by researchers at Stanford, UC Berkeley, and NVIDIA in June 2025, accepted at CoRL 2025, bringing the 'test-time scaling' idea from large language models over to VLAs. It leaves the original policy untouched: at deployment time, it samples several actions from the VLA, adds Gaussian perturbations, and uses majority voting to build a candidate set; a vision-language-model-based action verifier then scores the candidates and picks the best one to execute. The verifier is trained on automatically synthesized preference data (ranked by each candidate action's distance to the ground-truth action); the authors found that verification accuracy keeps improving with more synthetic data, and action error follows a roughly power-law relationship with the number of samples. Paired with models like OpenVLA, it delivers a 25-point absolute improvement on out-of-distribution tasks and 9 points on in-distribution tasks; an optimized serving engine can sample and verify 16 candidate actions in about 650 milliseconds.

ExampleWhen OpenVLA encounters an object it has never seen, RoboMonkey has it sample several sets of candidate grasping actions at once, and the verifier picks out the set most likely to succeed for execution.

Also called
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
Related
Inference-Time Compute · Best-of-N Sampling · Value-Guided Sampling · Vision-Language-Action Model · OpenVLA · Reward Model
Sources
arXiv 2506.17811: RoboMonkey
RoboMonkey 项目主页 (Chinese)
As of
2025-07

See it in the full glossary →