Action Representation
动作表示CommonThe quantity, reference frame, and format used to describe what action a robot should take.
Action representation is how a policy's output action gets written down, one of the first things to settle when building a robot-learning system. It involves several choices: what quantity to control (joint angle, end-effector pose, or velocity or torque instead); what it's relative to (a fixed global frame, an increment relative to the last step, or a whole segment relative to the current pose); how to write rotation (Euler angles, quaternions, or the 6D representation, which is more continuous and easier for a neural network to learn); and the output format (continuous values, discrete tokens, or a more abstract intermediate representation such as a latent action or trajectory points). Swapping one action representation for another, with the same model, can change success rate a great deal, so papers usually spell this choice out explicitly and compare alternatives.
ExampleThe UMI paper compares these on a cup-pouring task: a whole trajectory relative to the current end-effector pose succeeds 20 out of 20 times, step-by-step increments, whose errors accumulate, succeed 16 out of 20, and a global absolute-coordinate representation succeeds only 5 out of 20.
- Also called
- Action Parameterization
- Related
- Action Space · Delta (Relative) Action vs. Absolute Action · 6D Rotation Representation · Action Tokenizer · Latent Action · End-Effector Pose
- Sources
- Universal Manipulation Interface (UMI, arXiv:2402.10329)
On the Continuity of Rotation Representations in Neural Networks (arXiv:1812.07035)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective (arXiv:2507.01925)