Embodied AI Glossary中文

Temporal-Difference Learning

时序差分学习TDCommon

Updating a value estimate using “this step's reward plus the estimated value of the next state,” without waiting for the episode to end.

Temporal-difference (TD) learning was systematically introduced by Richard Sutton in a 1988 paper and is the core method for estimating a value function (how much total return follows from a given state or action) in reinforcement learning. It doesn't need to wait for an episode to finish to get the true return; instead it updates after every single step, treating “the reward r actually received, plus the discounted estimated value of the next state, γV(s′)” as a target, calling the gap between that target and the current estimate V(s) the TD error, and nudging V(s) toward the target by that amount. This “using an estimate to update an estimate” approach is called bootstrapping, and it has lower variance than Monte Carlo methods, which wait for the full episode, letting it learn while still acting — at the cost of introducing some bias. Q-learning, DQN, and the critic networks in SAC and TD3 are all trained with TD targets; the early landmark system TD-Gammon used it to reach expert-level backgammon play.

ExampleA robot arm takes one step and gets a reward of 0; the critic estimates the next state's value at 0.8, and with discount factor γ = 0.99, the TD target is 0 + 0.99 × 0.8 = 0.792. If the current state's estimated value is 0.5, the TD error is 0.292, and the network nudges its estimate toward 0.792.

Also called
TD Learning, TD Error
Related
Value Function · Bellman Equation · Bootstrapping (in Reinforcement Learning) · Q-Learning · Monte Carlo Methods · Generalized Advantage Estimation
Sources
Temporal difference learning (Wikipedia)
Soft Actor-Critic (OpenAI Spinning Up)

See it in the full glossary →