Embodied AI Glossary中文

AutoRT

Advanced

DeepMind's 2024 system that uses a large model to automatically assign tasks to a fleet of robots and collect real data.

AutoRT is a system Google DeepMind announced in January 2024, with authors including Karol Hausman, Fei Xia, Chelsea Finn, and Sergey Levine. Training embodied foundation models needs real data, but having one person watch one robot to collect it doesn't scale. AutoRT has an off-the-shelf large model act as dispatcher: robots explore autonomously in an office building, a vision-language model describes the scene and objects, and a large language model proposes tasks based on that. Each task is first filtered by a “robot constitution” (basic rules inspired by Asimov's Three Laws, plus safety rules and embodiment-capability rules), then assigned — depending on available staffing — to teleoperation, a scripted grasping policy, or RT-2 to execute. Over 7 months it orchestrated robots (up to more than 20 at once) across 4 buildings, collecting 77,000 real-robot episodes covering more than 6,650 distinct instructions, with one person able to supervise 3 to 5 robots at once.

ExampleIf the LLM proposes “open the fridge and get a drink,” the constitution's embodiment rule would filter it out for a single-arm robot, since the task needs two hands; tasks involving people or animals, sharp or fragile objects, or electrical appliances get blocked by the safety rules.

Also called
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
Related
Autonomous Data Collection · Asimov's Three Laws of Robotics · RT-2 · RT-1 · SayCan · LLM-based Task Planning
Sources
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents (arXiv 2401.12963)
Shaping the future of advanced robotics (Google DeepMind blog, 2024-01-04)
AutoRT 项目页 (Chinese)
As of
2024-01

See it in the full glossary →