Embodied Reasoning
具身推理ERCommonMaking a model understand the physical world: where things are, how to grasp them, and what to do next.
Embodied reasoning means a model's spatial, temporal, and causal understanding and inference about the physical world, put to use for robot action: pointing to a graspable location in an image, predicting an object's 3D bounding box and motion trajectory, judging whether a task has been completed, or breaking a long instruction into steps. When Google DeepMind released Gemini Robotics in March 2025, it packaged this capability into a separate model, Gemini Robotics-ER, covering object detection, pointing, trajectory prediction, grasp prediction, multi-view correspondence, and 3D bounding-box prediction, and open-sourced the 400-question ERQA evaluation set alongside it. As of 2026 the series has been updated to Gemini Robotics-ER 2, accessible through the Gemini API. It often serves as the “slow” system in a fast-slow dual-system setup, thinking things through first, then handing off action execution to a VLA model or a low-level controller.
ExampleGiven a photo of a kitchen and the instruction “put the cup in the sink,” an embodied reasoning model first marks the cup handle's location and a path of waypoints to the sink in the image, then hands this off to an action model to execute.
- Also called
- ER
- Related
- Embodied Reasoning Model · Gemini Robotics-ER · Spatial Reasoning · ERQA · NVIDIA Cosmos Reason · Dual-System Architecture (System 1 / System 2)
- Sources
- Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020)
ERQA benchmark (GitHub, Google DeepMind)
Gemini Robotics ER - Google DeepMind - As of
- 2026-09