Control Latency
控制延迟CommonThe time gap between a sensor capturing data and the motor actually executing the corresponding action.
Control latency is the total time a full sense-compute-execute cycle takes: camera exposure and data transfer, model inference, network or bus communication, and motor-drive response all count toward it, with end-to-end latency meaning the whole gap from when an observation is captured to when the corresponding action lands on the joints. Feedback control works by seeing an error and correcting it, so the larger the latency, the more outdated the information the controller is acting on — at best this makes motion sluggish and prone to overshoot, and at worst it causes oscillation and instability; the Smith predictor, proposed by O. J. M. Smith in 1957, is a classic technique for dealing with pure delay. A single VLA inference often takes tens to well over a hundred milliseconds, far longer than the underlying control cycle, which is why techniques like action chunking, asynchronous inference, and real-time chunking (RTC) exist. It's a broader notion than 'inference latency,' which measures only the model's own forward-pass time.
ExamplePhysical Intelligence's real-time chunking paper reports that π0.5 takes about 76 ms per inference on an RTX 4090, while the robot executes at 50 Hz (20 ms per step); by the time a new action is ready, the robot has already taken about 3 more steps, and without special handling, the boundary between action chunks shows up as a pause or a jump.
- Also called
- End-to-End Latency, Execution Latency, Perception-to-Action Latency
- Related
- Inference Latency · Control Frequency · Real-Time Chunking · Asynchronous Inference · Action Chunking · Jitter (Timing Jitter)
- Sources
- Real-Time Execution of Action Chunking Flow Policies (Black, Galliker, Levine, arXiv 2506.07339)
Wikipedia: Smith predictor - As of
- 2025-12