Embodied AI Glossary中文

Heterogeneous Pre-trained Transformers

异构预训练 TransformerHPTAdvanced

A 2024 method from Kaiming He's group at MIT that pretrains one shared backbone jointly across many different robots' data.

HPT was released in September 2024 by Lirui Wang and Kaiming He at MIT CSAIL together with Xinlei Chen at Meta FAIR, published at NeurIPS 2024 as a Spotlight paper. Robot data is highly “heterogeneous”: different robots vary in number of cameras, joint count, and control scheme, making it hard to train them all together inside one model. HPT splits the network into three parts: a small stem (encoder) per embodiment that converts proprioception and images into a fixed number of tokens; a large, shared Transformer trunk in the middle that learns a representation independent of embodiment and task; and an output-side head that maps that representation into an action for a given task. Pretraining used more than 50 datasets and about 200,000 trajectories, drawn from real-robot teleoperation, simulation, human video, and already-deployed robots.

ExampleAttaching the pretrained trunk to a new robot only requires training a new stem and head for it; the authors report more than 20% improvement in fine-tuned policy performance across several simulation benchmarks and unseen real-robot tasks.

Also called
HPT
Related
Cross-Embodiment · Heterogeneous Data · Pre-training · Embodiment-specific Head · Scaling Law · Open X-Embodiment
Sources
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers (arXiv 2409.20537)
HPT 项目页 (Chinese)
As of
2024-12

See it in the full glossary →