Embodied AI Glossary中文

Partially Observable Markov Decision Process

部分可观测马尔可夫决策过程POMDPAdvanced

The decision-making framework for when an agent can't see the true state, only noisy observations of it.

A Partially Observable Markov Decision Process extends the Markov Decision Process (MDP), the standard model for sequential decision-making, to cases where the agent cannot access the true state and only receives noisy, incomplete observations. Karl Åström introduced the framework in 1965, and Leslie Kaelbling, Michael Littman, and colleagues brought it systematically into AI planning in 1998. A common approach is to maintain a “belief” — a probability distribution over the true state — and update it with Bayes' rule each time a new observation arrives. Robots operate in partially observable settings almost by default: cameras get occluded, drawers hide their contents, and many objects only reveal what they are once picked up. Exact solutions are computationally intractable, so practical systems use approximate planning, or give the policy network a history of frames plus a memory module to implicitly estimate the state.

ExampleA robot searching for a cup in a row of closed cabinets can't see inside before opening a door, so it updates a probability distribution over “which compartment the cup is likely in” based on the compartments it has already checked, then decides which door to open next.

Also called
POMDP, Partially Observable MDP
Related
Markov Decision Process · Observation · State Space · Embodied Memory · Memory-Augmented VLA · Reinforcement Learning
Sources
Partially observable Markov decision process - Wikipedia

See it in the full glossary →