Embodied AI Glossary中文

Inference Latency

推理延迟Essential

The time a model takes from receiving input to producing output, which determines how quickly a robot can react.

Inference latency is the time a model takes, after receiving one frame of observation — image, joint state, instruction — to finish its forward computation and produce an action, usually measured in milliseconds. Large models are compute-heavy, and their latency often can't keep up with a robot's control frequency of tens to hundreds of hertz: the 7-billion-parameter OpenVLA produces an action at only about 6Hz on an RTX 4090; the π0 paper measured about 73 milliseconds for one full inference with 3 camera views on the same GPU, or about 86 milliseconds when computed on a separate computer and sent over Wi-Fi. Too much latency makes a robot pause between action segments, or fail to keep up with a moving object. Common fixes include action chunking, outputting dozens of steps at once; asynchronous inference, computing the next segment while the current one executes; quantization and inference acceleration; and splitting large and small models into layers that each run at their own frequency.

Exampleπ0 controls a 50Hz robot by outputting 50 steps of action per inference call, executing 25 of them, 0.5 seconds, before running inference again for the next segment, instead of running the large model at every single step.

Also called
Inference Delay, Model Latency
Related
Inference · Action Chunking · Asynchronous Inference · Real-Time Chunking · Control Frequency · Post-Training Quantization
Sources
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)
Real-Time Execution of Action Chunking Flow Policies (arXiv:2506.07339)
As of
2025-06

See it in the full glossary →