Embodied AI Glossary中文

Heterogeneous Data

异构数据Advanced

Robot data from different robots, sensors, and collection methods, with mismatched formats and meanings.

“Heterogeneity” in robot data shows up on several levels at once: different embodiments (single-arm, bimanual, humanoid, each with its own joint count and action space); different sensors (number and placement of cameras, presence or absence of depth or touch); different control schemes (joint angles versus end-effector pose, different control frequencies); and different sources (real-robot teleoperation, simulation, human video). Any single lab's own data is too small on its own to train a general-purpose policy, so labs need to mix this kind of data together — but simply concatenating it confuses a model about what the same number means on different robots. A common fix gives each embodiment its own input encoder and output head, while sharing one large backbone in the middle — for instance, Kaiming He's group used this structure in HPT to pool 52 datasets together. Other approaches unify actions into one shared space, or adjust the data mixing ratio by source. Open X-Embodiment, which pools data from 22 kinds of robots, is a textbook example of a heterogeneous dataset.

ExampleHPT combines 52 datasets — real robots, multiple simulators, and human video — into one pretraining run; each embodiment has its own “stem” that compresses vision and proprioception into a small number of tokens, which are then fed into a shared Transformer backbone.

Also called
Multi-Source Heterogeneous Data
Related
Cross-Embodiment Data · Cross-Embodiment · Heterogeneous Pre-trained Transformers · Unified Action Space · Embodiment-specific Head · Data Mixture
Sources
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers (arXiv)
Open X-Embodiment: Robotic Learning Datasets and RT-X Models (arXiv)

See it in the full glossary →