Action Horizon
动作视界CommonHow many steps of history a policy looks at, how many steps of action it predicts, and how many it actually executes.
Action horizon describes how far a policy looks, in time, in both directions. The Diffusion Policy (2023) paper splits it into three quantities: the observation horizon To, how many recent steps of observation are fed in; the prediction horizon Tp, how many steps of action get generated in one shot; and the execution horizon Ta, how many of those actually get sent to the robot before replanning, called receding-horizon control. After executing Ta steps, the policy re-observes and predicts again. A longer Ta means smoother, more coherent motion and fewer inference calls, but a slower reaction to sudden changes; a shorter Ta is more responsive but can jitter. The paper found 8 steps worked best for most tasks, and the open-source code defaults to To=2, Tp=16, Ta=8. What VLAs call chunk size roughly corresponds to the prediction horizon.
Exampleπ0 predicts 50 steps of action at once: on a 50Hz robot, it executes only the first 25 steps, 0.5 seconds, before re-running inference with new images, while on a 20Hz UR5e it re-infers every 16 steps.
- Also called
- Prediction Horizon, Execution Horizon, Observation Horizon
- Related
- Action Chunking · Temporal Ensembling · Diffusion Policy · Asynchronous Inference · Closed-loop Control · Real-Time Chunking
- Sources
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv:2303.04137)
diffusion_policy 训练配置 train_diffusion_unet_image_workspace.yaml (GitHub) (Chinese)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)