Embodied AI Glossary中文

Monte Carlo Methods

蒙特卡洛方法(蒙特卡洛回报)MCAdvanced

Estimating a value by averaging over many random samples; in RL, using a whole episode's actual return to estimate value.

Monte Carlo methods broadly refers to any technique that approximates a computation through random sampling and averaging. In reinforcement learning it refers specifically to this: let the agent run a full episode to completion, sum up the actual (discounted) rewards received from some step onward to get the “Monte Carlo return,” then average that over many episodes to estimate the value of a state or action. It does not rely on any estimate of later states' value, so it avoids the bias that bootstrapping introduces, but variance grows with episode length, and it can only update once an episode ends, so it only applies to tasks that actually terminate. Its counterpart is temporal difference learning, which can update at every single step; TD(λ) lets you tune continuously between the two. Monte Carlo tree search borrows the same idea of random simulation.

Exampleπ*0.6's RECAP training uses the Monte Carlo return directly for its value function: each step costs a reward of −1, a successful finish scores 0, and a failure subtracts a large penalty, so the value roughly equals the negative of how many steps remain to complete the task.

Also called
MC, Monte Carlo Return
Related
Return · Temporal-Difference Learning · Value Function · Discount Factor · Monte Carlo Tree Search · RECAP
Sources
Wikipedia: Reinforcement learning (Monte Carlo methods)
Physical Intelligence 2025: π*0.6: a VLA That Learns From Experience (RECAP)

See it in the full glossary →