RT-1
EssentialGoogle's 2022 Transformer policy for real robots, trained on 130,000-plus real-robot demonstrations across 700-plus tasks.
RT-1 is a robot control model released by the Google Robotics team and Everyday Robots in December 2022. At the time, most robot learning trained one small model per task; RT-1 set out to test whether the recipe behind language and vision breakthroughs — a large-capacity model trained on large, diverse data — would also work for robots. The team used 13 mobile manipulators over 17 months to collect more than 130,000 demonstrations across more than 700 tasks. The model takes in images and a language instruction: images go through an EfficientNet feature extractor, FiLM layers (which let the language instruction modulate the visual features) fold in the instruction, TokenLearner compresses the number of tokens, and a Transformer outputs discretized actions, controlling the arm and base in closed loop at 3Hz. It reached 97% success on trained instructions and generalized to new tasks, distractors, and backgrounds better than the baselines of the time. Its data later became part of Open X-Embodiment, and both its architecture and data underpin RT-2.
ExampleIn a Google office kitchen environment, RT-1 followed language instructions to pick and place objects, open and close drawers, and put items away in drawers.
- Also called
- Robotics Transformer 1, RT-1: Robotics Transformer for Real-World Control at Scale
- Related
- RT-2 · RT-X · RT-1 Robot Action Dataset · Transformer · EfficientNet · TokenLearner
- Sources
- RT-1: Robotics Transformer for Real-World Control at Scale (arXiv 2212.06817)
RT-1 项目主页 (Chinese) - As of
- 2022-12