Credit Assignment
信用分配AdvancedFiguring out, once a final reward or penalty arrives, which of the earlier actions deserve the credit or the blame.
The credit assignment problem was already discussed specifically in Marvin Minsky's 1961 survey “Steps Toward Artificial Intelligence”: once a complex strategy succeeds, how should the credit be divided among the many decisions involved? His example was that winning a game of chess might involve a million decisions, and asked whether each one could just get an equal millionth of the credit. In reinforcement learning it mainly refers to credit assignment across time: reward is often delayed until a task finishes, mixed in with noise and chance along the way, and the agent has to learn each action's actual contribution to the final outcome from limited experience. Common tools include temporal-difference learning with bootstrapping, the discount factor, the advantage function and Generalized Advantage Estimation (GAE), reward shaping (adding intermediate reward), and hierarchical reinforcement learning. Robot long-horizon tasks have many action steps and sparse reward, which makes credit assignment one of the main reasons reinforcement learning is hard to apply there.
ExampleA robot arm stacks 5 blocks and gets 1 point only if all 5 end up stacked correctly. If one episode fails, the cause might be block 2 placed crooked, or block 4 released too early; credit assignment has to learn how much blame each step's action deserves from a large number of episodes like this.
- Also called
- Credit Assignment Problem, CAP
- Related
- Sparse Reward · Temporal-Difference Learning · Advantage Function · Generalized Advantage Estimation · Reward Shaping · Long-horizon Task
- Sources
- A Survey of Temporal Credit Assignment in Deep Reinforcement Learning (arXiv 2312.01072)
Minsky (1961): Steps Toward Artificial Intelligence (Proceedings of the IRE)