Observation
观测EssentialThe information an agent receives from the environment at each moment, the input a policy uses to decide.
An observation is the input an agent receives from the environment at each step. In reinforcement learning's observation–action–reward loop, the environment returns a new observation both on reset and after every action; the observation space defines the format and range of these inputs — in Gymnasium, for example, CartPole's observation is just a few numbers: cart position, velocity, pole angle, and so on. A real robot's observation typically includes several camera images, depth or point-cloud data, proprioceptive information such as joint angles and gripper opening, and a language instruction. Observation is not the same as state: state is the environment's complete description, while an observation often only reveals part of it, with occlusion or noise — this situation is modeled as a Partially Observable Markov Decision Process (POMDP). Many policies take in the last several frames of observation; Diffusion Policy calls this window the observation horizon.
ExampleA VLA policy's observation at each step: one RGB image each from a head camera and a wrist camera, seven joint angles, gripper opening, plus the instruction “put the cup on the plate.”
- Also called
- Observation Space
- Related
- State Space · Action Space · Proprioception · Partially Observable Markov Decision Process · Policy · Observation-Action Pair
- Sources
- Basic Usage (Gymnasium Documentation)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv:2303.04137)