Embodied AI Glossary中文

Inference-Time Compute

推理时计算Advanced

Improving results by spending more compute at inference time, instead of changing the trained model.

Inference-time compute means spending more computation at inference time, after training is already finished, to improve results: generating a longer chain of thought, sampling many candidate answers and using a verifier (a model that scores candidates) to pick the best one, or repeatedly revising an answer. OpenAI's o1 (2024) and a paper by Snell and colleagues brought wide attention to this direction; the latter found that allocating inference compute according to a problem's difficulty let a small model outperform one 14 times larger on some problems. It offers a path to better performance besides simply making the model bigger. Robotics has an analogous technique: sampling a VLA's actions multiple times and using a verifier to select among them, at the cost of higher inference latency.

ExampleRoboMonkey (2025) has a VLA sample multiple actions for the same observation, adds Gaussian perturbations, takes a majority vote, and then uses a VLM verifier to pick the best action; the paper reports about a 25-percentage-point absolute improvement in success rate on out-of-distribution tasks.

Also called
Test-Time Scaling, Test-Time Compute
Related
Best-of-N Sampling · Value-Guided Sampling · RoboMonkey · Chain-of-Thought · Inference Latency · Scaling Law
Sources
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (arXiv:2408.03314)
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models (arXiv:2506.17811)
As of
2025-07

See it in the full glossary →