Embodied AI Glossary中文

RT-X

Common

RT-1 and RT-2 retrained on Open X-Embodiment, a large dataset pooling robot data across 22 different robot embodiments.

RT-X is a set of models released in October 2023, led by Google DeepMind in collaboration with 21 institutions and 34 labs, alongside the Open X-Embodiment dataset. That dataset pools 60 existing robot datasets, 22 different robot embodiments, and more than 1 million real-robot trajectories. RT-1-X and RT-2-X keep the architectures of RT-1 (a dedicated robot Transformer) and RT-2 (a VLA fine-tuned from a vision-language model) unchanged, simply swapping in this combined dataset for training. The results show positive transfer from cross-embodiment training: RT-1-X beat the original methods used by individual labs by about 50% on average in low-data settings, and RT-2-X scored roughly 3x higher than RT-2 on emergent-skill evaluations. It drove wider cross-embodiment data sharing, and later models such as Octo and OpenVLA are all trained on this same dataset.

ExampleRT-2-X picked up spatial concepts that were only present in other labs' data, learning to tell apart instructions that differ by a single preposition, such as “put the apple on the cloth” versus “put the apple next to the cloth.”

Also called
RT-1-X, RT-2-X
Related
Open X-Embodiment · RT-1 · RT-2 · Cross-Embodiment · Octo · OpenVLA
Sources
Open X-Embodiment: Robotic Learning Datasets and RT-X Models 项目主页 (Chinese)
As of
2023-10

See it in the full glossary →