Gemini Robotics-ER
CommonGoogle DeepMind's embodied-reasoning model, which understands spatial layouts, makes plans, and judges whether a robot task is done.
ER stands for Embodied Reasoning. Google DeepMind introduced it in March 2025 alongside the original Gemini Robotics: a vision-language model built on Gemini with strengthened 3D perception and spatial understanding, able to point out object locations and predict grasp points and trajectories. It generally doesn't output joint-level actions directly; instead it answers questions like where something is, which step comes first, and whether a step is complete, then hands the plan to a VLA or robot interface to execute — functioning as the “slow” system in a fast-slow dual-system setup. ER 1.5, from September 2025, is available through the Gemini API, later updated to 1.6; ER 2, from July 2026, adds video understanding and progress tracking, and can orchestrate action models through multistep tasks, coordinate multiple robots, and call tools such as Google Search.
ExampleER 2 watches a video of a robot at work, judges whether the current step is finished, has the robot retry if not, and moves on to the next step once it is.
- Also called
- Gemini Robotics-ER 1.5, Gemini Robotics-ER 1.6, Gemini Robotics ER 2
- Related
- Embodied Reasoning · Embodied Reasoning Model · Gemini Robotics · Gemini Robotics 1.5 · ERQA · Brain–Cerebellum Architecture
- Sources
- Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020)
Gemini Robotics ER - Google DeepMind
Gemini Robotics ER 2 (Google blog) - As of
- 2026-07