Embodied AI Glossary中文

Observation-Action Pair

观测-动作对Common

One training sample: what the robot 'saw' at a given moment paired with what it 'did' next.

This is the basic unit of imitation-learning data. At every timestep, an observation o (camera images, joint angles and other proprioceptive state, often with a language instruction attached) and the action a executed at that moment (target joint angles, end-effector displacement, gripper opening) are recorded together; a full demonstration trajectory is a time-ordered sequence of these (o, a) pairs. Behavior cloning treats this as supervised learning: given o, predict a. Strictly speaking, “state” is a complete description of the environment, while “observation” is only the part sensors can actually see, and on a real robot usually only observations are available. Because these samples are sequentially correlated, once a policy drifts even slightly it runs into observations it never saw in training, and the error compounds — this is called compounding error. In a LeRobot dataset, each frame's observation.state, observation.images, and action fields are exactly this structure.

ExampleRecording a demonstration of an SO-101 arm grasping a block at 30 frames per second in LeRobot, every frame stores the camera image, the follower arm's joint readings, and the leader arm's target joint angles at that moment as the action.

Also called
State-Action Pair
Related
Observation · Action Label · Behavior Cloning · Trajectory · Demonstration Data · Compounding Error
Sources
LeRobotDataset v3.0 (Hugging Face LeRobot docs)
google-research/rlds (GitHub)

See it in the full glossary →