Embodied AI Glossary中文

Action Smoothing

动作平滑Common

Filtering or penalizing a policy's output actions to remove high-frequency jitter, so joint motion stays smooth.

Action smoothing is a common step when deploying a learned policy. A neural network outputs an action independently at every step, so adjacent steps can jump around, making motors jitter back and forth, waste power, and heat up. There are three common fixes. One is filtering the action before sending it out, letting only low-frequency changes through — the simplest is an exponential moving average, y_t = α·x_t + (1−α)·y_{t−1}, where x is the policy's newly output action and y is the action actually sent, with smaller α giving more smoothing. Another is adding a penalty during training, such as the action_rate term in legged_gym's default reward, which penalizes the difference between consecutive actions; the CAPS method (ICRA 2021) adds both temporal and spatial smoothness regularization, and its authors report nearly an 80% reduction in power use on a quadrotor. A third is predicting a whole chunk of actions at once (action chunking), then temporally ensembling or interpolating them. The cost is that filtering introduces delay (phase lag), making the robot slower to react to disturbances, so α has to be tuned as a tradeoff between smoothness and responsiveness.

ExampleA quadruped's reinforcement-learning policy outputs target joint angles at 50 Hz; deployed on the real robot, its legs show noticeable high-frequency jitter. Adding a first-order low-pass filter noticeably reduces the jitter, but the robot also reacts a bit more slowly when pushed, requiring α to be re-tuned.

Also called
Low-Pass Filtering, Action Filtering, Action Smoothness Regularization
Related
Action Chunking · Temporal Ensembling · Control Frequency · Jerk · Real-Time Chunking · Proportional-Derivative Control
Sources
Regularizing Action Policies for Smooth Control with Reinforcement Learning (CAPS, arXiv 2012.06644)
Wikipedia: Low-pass filter
legged_gym legged_robot_config.py(action_rate 奖励项) (Chinese)

See it in the full glossary →