Embodied AI Glossary中文

Soft Actor-Critic

软演员-评论家SACCommon

An off-policy reinforcement-learning algorithm that pursues high return while also encouraging the policy to stay somewhat random.

SAC was introduced by UC Berkeley's Tuomas Haarnoja, Sergey Levine, and colleagues in 2018 (ICML 2018). It's an actor-critic method: the “actor” is the policy network that outputs actions, and the “critic” is a Q-function estimating how good an action is. Its core idea is the maximum-entropy objective: maximize return while also keeping the policy's entropy — how random it is — as high as reasonably possible, which encourages exploration and avoids collapsing onto a suboptimal solution too early; the degree of randomness is set by a temperature coefficient α, which later versions can tune automatically. It's off-policy, able to reuse old data from an experience-replay buffer repeatedly, which gives it good sample efficiency and low sensitivity to hyperparameters, though it applies only to continuous action spaces. Implementation-wise, it trains two Q-networks and takes the smaller estimate, to counter Q-value overestimation. SAC is a common foundation for real-robot reinforcement learning; the RLPD algorithm used by SERL and HIL-SERL is itself a refinement built on SAC.

ExampleThe SAC extension paper applies it directly on real hardware: a Minitaur quadruped learns to walk in about 2 hours, a Sawyer arm learns to stack blocks in about 2 hours, and a dexterous hand learns to turn a valve directly from images in about 20 hours.

Also called
SAC, Maximum Entropy Actor-Critic
Related
Off-Policy · Entropy Regularization · Q-Function · Experience Replay · Twin Delayed DDPG · SERL
Sources
Soft Actor-Critic: Off-Policy Maximum Entropy Deep RL with a Stochastic Actor (arXiv 1801.01290)
Soft Actor Critic—Deep Reinforcement Learning with Real-World Robots (BAIR Blog, 2018)
Soft Actor-Critic (OpenAI Spinning Up)

See it in the full glossary →