Embodied AI Glossary中文

Delta (Relative) Action vs. Absolute Action

增量动作 / 绝对动作Common

Writing an action as “move by this much” (delta) versus “go to this coordinate” (absolute).

This is one of the basic choices in action representation. An absolute action gives a target value directly, such as the pose the end-effector, the gripper or tool at the very tip of the arm, should reach, or a target angle for each joint. A delta action instead gives only the change relative to the current state, such as “move 1 centimeter further along x.” Delta actions don't depend on precise calibration between the robot's base and the world frame, and their value distribution is more consistent across different scenes, but errors accumulate as they add up step by step. Absolute actions don't accumulate error, but they need the coordinate frame to be well calibrated. There's a middle ground too: UMI uses a “relative trajectory,” writing a whole action segment as a transform relative to the end-effector pose at the start of that segment. The Diffusion Policy paper separately found that, for its test tasks, position control that outputs a target position beat velocity control that outputs a target velocity. Which one to use depends on how the dataset defines it, and training and deployment must match, or the robot will drift off course.

ExampleIf the arm's end effector is currently at x=0.40 meters and needs to move to x=0.42 meters, an absolute action writes 0.42, and a delta action writes +0.02. In one UMI experiment, an absolute-action baseline succeeded only 25% of the time because SLAM coordinates and the robot base's coordinates were hard to align, while the relative-trajectory version succeeded 100% of the time.

Also called
Relative Action, Delta Action, Absolute Action, Relative Trajectory
Related
Action Space · Action Representation · End-Effector Pose · Coordinate Frame · Universal Manipulation Interface · Action Chunking
Sources
Universal Manipulation Interface (UMI, arXiv 2402.10329)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)

See it in the full glossary →