Embodied AI Glossary中文

Deep Deterministic Policy Gradient

深度确定性策略梯度DDPGAdvanced

An actor-critic algorithm that extends DQN to continuous actions, with the policy outputting one deterministic action directly.

DDPG was introduced by DeepMind's Lillicrap and colleagues in 2015 (ICLR 2016), building on the deterministic policy gradient (DPG) theory from Silver and colleagues (2014). DQN can only handle discrete actions, because it needs to search over every action for the one with the highest Q-value, and continuous actions like a robot arm's joint angles can't be enumerated one by one. DDPG uses two networks: a critic that learns the Q-function, and an actor (the policy network) that outputs one deterministic action directly, updated along the gradient direction that increases the Q-value. It's off-policy, reusing DQN's experience replay and target network, and adds noise to the action during training for exploration. The original paper validated it on more than 20 simulated physics tasks, several of which could be learned directly from pixels. DDPG is sensitive to hyperparameters and prone to overestimating Q-values; the later TD3 (Twin Delayed DDPG) specifically fixes the overestimation problem, and algorithms like SAC are more commonly used in practice today.

ExampleOpenAI's 2017 hindsight experience replay (HER) experiments used DDPG to train a 7-DOF Fetch arm in simulation to push, slide, and pick-and-place objects, and deployed the resulting policy to a real robot.

Also called
DDPG
Related
Twin Delayed DDPG · Soft Actor-Critic · Deep Q-Network · Deterministic vs. Stochastic Policy · Off-Policy · Hindsight Experience Replay
Sources
Continuous control with deep reinforcement learning (arXiv 1509.02971)
OpenAI Spinning Up: Deep Deterministic Policy Gradient
Hindsight Experience Replay (arXiv 1707.01495)

See it in the full glossary →