Unified Action Space
统一动作空间CommonA single fixed format that maps different robots' actions by physical meaning, so their data can be trained together.
Different robots use different action formats: a single arm might be 7 dimensions, a bimanual robot 14; some give joint angles, others give end-effector pose (the gripper's position and orientation in space). To train one model on data from multiple robots, the actions first need to be aligned into a common format — this is a unified action space. The most direct approach is to define a vector long enough for everything, with each position fixed to a physical quantity and missing dimensions padded with zero: Tsinghua's RDT-1B uses 128 dimensions, with fixed slots for the left arm, right arm, and base; π0 zero-pads every robot's action to 18 dimensions. A different approach learns an abstract latent action space instead, as in UniAct, then adds a small per-robot decoder to turn it back into a concrete command. This is a foundational piece of cross-embodiment training.
ExampleFitting data from a 6-DOF single-arm robot into RDT-1B's 128-dimensional vector, the arm is treated as the right arm: its joint angles fill only the first 6 slots of the right-arm block, while the left-arm and base slots are all set to zero.
- Also called
- Universal Action Space
- Related
- Cross-Embodiment · Action Space · Embodiment-specific Head · Latent Action · Cross-Embodiment Data · RDT-1B
- Sources
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation (arXiv 2410.07864)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)
Universal Actions for Enhanced Embodied Foundation Models (arXiv 2501.10105)