Embodied AI Glossary中文

Gemini Robotics-ER

Common

Google DeepMind's embodied-reasoning model, which understands spatial layouts, makes plans, and judges whether a robot task is done.

ER stands for Embodied Reasoning. Google DeepMind introduced it in March 2025 alongside the original Gemini Robotics: a vision-language model built on Gemini with strengthened 3D perception and spatial understanding, able to point out object locations and predict grasp points and trajectories. It generally doesn't output joint-level actions directly; instead it answers questions like where something is, which step comes first, and whether a step is complete, then hands the plan to a VLA or robot interface to execute — functioning as the “slow” system in a fast-slow dual-system setup. ER 1.5, from September 2025, is available through the Gemini API, later updated to 1.6; ER 2, from July 2026, adds video understanding and progress tracking, and can orchestrate action models through multistep tasks, coordinate multiple robots, and call tools such as Google Search.

ExampleER 2 watches a video of a robot at work, judges whether the current step is finished, has the robot retry if not, and moves on to the next step once it is.

Also called
Gemini Robotics-ER 1.5, Gemini Robotics-ER 1.6, Gemini Robotics ER 2
Related
Embodied Reasoning · Embodied Reasoning Model · Gemini Robotics · Gemini Robotics 1.5 · ERQA · Brain–Cerebellum Architecture
Sources
Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020)
Gemini Robotics ER - Google DeepMind
Gemini Robotics ER 2 (Google blog)
As of
2026-07

See it in the full glossary →