Distributional Value Function
分布式价值函数AdvancedA value function that predicts the full probability distribution of future return, not just its average.
An ordinary value function outputs only the expected value of future return (cumulative reward); a distributional value function outputs the entire distribution of possible return values instead. DeepMind's Bellemare, Dabney, and Munos systematically introduced this view in a 2017 ICML paper, giving a distributional Bellman equation and the C51 algorithm (which represents the return distribution with 51 discrete support points), reaching state-of-the-art results on Atari at the time. A common implementation splits the return into several bins, has the network output a probability for each bin, and trains with cross-entropy, which is more stable than regressing a single number directly and also captures uncertainty. Recent VLA reinforcement-learning methods commonly use this as the critic.
ExamplePhysical Intelligence's π*0.6 (RECAP) trains a multi-task distributional value function: it discretizes the “steps remaining until success” return into bins and trains with cross-entropy, then uses it to estimate advantages for filtering good actions; AgiBot's LWD (2026) similarly uses distributional implicit value learning (DIVL) to handle sparse-reward data collected by a robot fleet.
- Also called
- Distributional Reinforcement Learning, Distributional RL, Value Distribution
- Related
- Value Function · Bellman Equation · Return · RECAP · π*0.6 · Cross-Entropy
- Sources
- A Distributional Perspective on Reinforcement Learning (arXiv:1707.06887)
π*0.6: a VLA That Learns From Experience (arXiv:2511.14759)
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies (arXiv:2605.00416) - As of
- 2026-09