RDT-1B
RDTCommonTsinghua's 2024 diffusion foundation model for bimanual manipulation, with about 1.2 billion parameters and a unified action space.
RDT-1B (Robotics Diffusion Transformer) is a robot diffusion foundation model, about 1.2 billion parameters, released by a Tsinghua University team in October 2024. In bimanual manipulation the same situation often has several equally valid ways to act (action multimodality), and direct regression tends to average them into a wrong action; RDT instead uses a diffusion Transformer that denoises step by step to generate an action sequence, which can represent this kind of distribution. To make use of data from many different robots, it designs a “physically interpretable unified action space” that places physical quantities like joint angles and end-effector pose from different robots into fixed positions in one shared vector. The model is first pretrained on more than 1 million trajectories across 46 datasets, then fine-tuned on more than 6,000 ALOHA bimanual demonstrations, after which it can learn a new skill from just 1 to 5 demonstrations. Its successor is RDT2.
ExampleGiven only 1 to 5 demonstrations, RDT-1B can learn a new bimanual manipulation skill on an ALOHA robot, and it can even execute zero-shot on objects and scenes it has never seen.
- Also called
- RDT, Robotics Diffusion Transformer
- Related
- Diffusion Policy · Diffusion Transformer · Bimanual Manipulation · Unified Action Space · Action Multimodality · RDT2
- Sources
- RDT-1B 项目主页 (Chinese)
- As of
- 2024-10