CrossFormer
AdvancedA single set of weights that controls single arms, bimanual arms, quadrupeds, ground vehicles, and drones with one cross-embodiment Transformer.
CrossFormer was proposed in August 2024 by Sergey Levine's group at Berkeley with CMU, an oral presentation at CoRL 2024. Different robots vary in camera count, proprioception, action dimensionality, and control frequency, and earlier cross-embodiment training often had to hand-align observation and action spaces, or drop some inputs. CrossFormer instead cuts multi-camera images, proprioception, and the task (language or a goal image) all into tokens arranged in a sequence, fed into a single decoder-only Transformer shared across all embodiments; readout tokens are inserted into the sequence, then routed to different action heads by embodiment category, each outputting an action chunk of the matching dimensionality — such as a 7D end-effector delta for a single arm, 14D joint positions for a bimanual robot, or a 2D waypoint for navigation. The model, about 130 million parameters, is trained on 900,000 trajectories across 20 embodiments; on real robots it matches embodiment-specific policies and clearly beats earlier cross-embodiment methods.
ExampleThe same CrossFormer outputs a 4-step, 7-dimensional end-effector action chunk when connected to a single-arm WidowX, a 100-step, 14-dimensional joint-position chunk when connected to a bimanual ALOHA, and a 12-dimensional joint target at every step when connected to a Go1 quadruped.
- Also called
- Scaling Cross-Embodied Learning, CrossFormer: Scaling Cross-Embodied Learning (One Policy for Manipulation, Navigation, Locomotion and Aviation)
- Related
- Cross-Embodiment · Octo · Embodiment-specific Head · Open X-Embodiment · Heterogeneous Pre-trained Transformers · Action Chunking
- Sources
- Scaling Cross-Embodied Learning / CrossFormer (arXiv:2408.11812)
CrossFormer 项目主页(CoRL 2024 Oral) (Chinese) - As of
- 2024-08