Embodied AI Glossary中文

Asynchronous Inference

异步推理Common

The robot keeps executing the current action chunk while the model computes the next one, with no pause between.

Asynchronous inference splits “the model computing an action” and “the robot executing an action” into two parallel processes. In synchronous inference, the robot has to stop and wait once it finishes executing an action chunk until the model computes the next one, and the pause grows more noticeable the bigger the model. The asynchronous approach instead sends the latest observation to the model once the action queue drops below some threshold, so the next chunk finishes computing before the current one runs out, with the overlapping portion merged according to some rule. Hugging Face introduced this mechanism alongside SmolVLA in June 2025, splitting LeRobot into a PolicyServer that runs the model and a RobotClient that controls the robot. The hard part is stitching chunks together: a new chunk is computed from a slightly earlier observation, so switching directly can cause a jump in motion, which is exactly what Physical Intelligence's real-time chunking (RTC) was designed to handle.

ExampleIn LeRobot, with each inference outputting 50 steps of action and a threshold of 0.5, once fewer than 25 steps remain in the queue, the client sends a new image to the server, and the overlapping portion is merged with a weighted average.

Also called
Async Inference
Related
Inference Latency · Action Chunking · Real-Time Chunking · Policy Server (Remote Inference) · SmolVLA · LeRobot
Sources
Asynchronous Inference (LeRobot 文档) (Chinese)
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics (arXiv:2506.01844)
Real-Time Execution of Action Chunking Flow Policies (arXiv:2506.07339)
As of
2025-06

See it in the full glossary →