RoboAgent
RoboAgent(MT-ACT)MT-ACTAdvancedA multi-skill kitchen manipulation agent from CMU and Meta, trained on just 7,500 demonstrations.
RoboAgent was released by Carnegie Mellon University and Meta AI in September 2023 (accepted at ICRA 2024), addressing the question of whether a small amount of real-robot data can still train a robot with many skills, given how expensive that data is. The approach combines two ideas. First, semantic augmentation: the Segment Anything Model (SAM) segments objects in each frame, and their shape, color, and texture are then altered to multiply the existing data. Second, MT-ACT (Multi-Task Action Chunking Transformer), which extends ACT's action chunking (predicting a short block of future actions at once) to a multi-task setting where a language instruction distinguishes between tasks. Using only 7,500 teleoperated trajectories, a single policy learned 12 skills across 38 kitchen tasks, outperforming prior methods by more than 40% in unseen scenarios. The associated data was open-sourced under the name RoboSet.
ExampleA single demonstration of 'pull open the drawer' can be semantically augmented into many training samples with different-looking drawers and countertops, all used together to train MT-ACT.
- Also called
- MT-ACT, Multi-Task Action Chunking Transformer, RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking
- Related
- Action Chunking · Action Chunking with Transformers · Data Augmentation · Segment Anything Model · Multi-Task Learning · Language-conditioned Policy
- Sources
- arXiv 2309.01918: RoboAgent
RoboAgent 项目主页(RoboPen) (Chinese) - As of
- 2024-05