Gemini Robotics
EssentialGoogle DeepMind's 2025 VLA built on Gemini 2.0, able to directly control robots for dexterous manipulation.
Gemini Robotics is a vision-language-action model released by Google DeepMind on March 12, 2025, built on top of the Gemini 2.0 multimodal model. Released alongside it, Gemini Robotics-ER focuses on spatial understanding and embodied reasoning, such as object detection and predicting trajectories and grasps. The goal is to bring the general understanding of large models into the physical world, and Google highlighted three properties: generalization (handling new objects, scenes, and instructions, scoring more than twice as well as other leading VLAs at the time on a general-generalization benchmark), interactivity (understanding conversational instructions and adjusting in real time as the environment or instructions change), and dexterity (such as folding origami or packing a snack into a zip-lock bag). It was trained mainly on the bimanual ALOHA 2 platform, but also adapts to Franka arms and Apptronik's Apollo humanoid robot; the technical report says a new task can be fine-tuned with around 100 demonstrations. Later versions include On-Device and 1.5.
ExampleIn official demos, an ALOHA 2 bimanual robot running Gemini Robotics folds origami and packs a snack into a zip-lock bag on spoken command; if someone moves the target object mid-task, the robot adjusts its motion and keeps going.
- Also called
- Gemini Robotics 1.0
- Related
- Gemini Robotics-ER · Gemini Robotics 1.5 · Gemini Robotics On-Device · Vision-Language-Action Model · ALOHA 2 · Google DeepMind
- Sources
- Gemini Robotics brings AI into the physical world (Google DeepMind blog)
Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020) - As of
- 2025-03