Embodied AI Glossary中文

Return

回报Common

The sum of all rewards from a given moment onward, usually discounted so that later rewards count for less.

Return is the quantity reinforcement learning aims to maximize. A reward is the immediate score the environment gives at each step; the return is the total of all rewards from the current moment onward, often written G_t or R(τ). When an episode has a fixed length, these can just be added up directly; for long or never-ending tasks, a discounted return is normally used instead: a reward k steps in the future is multiplied by the discount factor γ raised to the k-th power (γ is between 0 and 1, commonly 0.95–0.99), so more distant rewards count for less — this both keeps the sum finite and makes the agent weigh near-term outcomes more. Reinforcement learning's objective is to maximize expected return; the value function and Q-function are exactly estimates of expected future return. Methods like return conditioning even feed a target return in as an input, so the policy acts at a specified performance level.

ExampleA grasping task gives +1 only on success and 0 otherwise, with γ = 0.99: succeeding at step 10 gives a discounted return from the start of about 0.99^10 ≈ 0.904, while dragging it out to step 50 gives only about 0.605 — so the policy is pushed toward finishing faster.

Also called
Cumulative Reward, Discounted Return, G_t
Related
Reward Function · Discount Factor · Value Function · Q-Function · Episode · Return Conditioning
Sources
OpenAI Spinning Up: Key Concepts in RL
Hugging Face Deep RL Course: The Reinforcement Learning Framework

See it in the full glossary →