Markov Decision Process
马尔可夫决策过程MDPCommonThe standard mathematical framework for an agent's loop of seeing a state, acting, getting a reward, and moving to a new state.
The Markov decision process is reinforcement learning's standard mathematical model, usually written as the tuple (S, A, P, R, γ): the state space, the action space, the state-transition probability (the chance of reaching a given next state after taking action a in state s), the reward function, and a discount factor that discounts distant rewards. Its core assumption is the Markov property: the next state depends only on the current state and action, not on earlier history. Its mathematical foundation comes from Richard Bellman's dynamic-programming work around 1957, and the Bellman equation, value functions, and policy gradients are all built on top of it. Reinforcement learning's goal is to find a policy within an MDP that maximizes expected cumulative discounted return. Real robots often can't observe the full state (an object might be occluded), which calls for a partially observable Markov decision process (POMDP) instead; in practice this is often handled by feeding in a history of past observations.
ExampleWhen training a quadruped to walk, the state can include joint angles, joint velocities, body orientation, and the commanded velocity; the action is the 12 joints' target angles; the reward rewards tracking the commanded velocity and penalizes energy use and falling; and the policy outputs one action per control cycle based on the current state.
- Also called
- MDP
- Related
- Reinforcement Learning · Partially Observable Markov Decision Process · Bellman Equation · Reward Function · Discount Factor · Policy
- Sources
- Markov decision process(Wikipedia)