Embodied AI Glossary中文

Gemini Robotics 1.5

Common

Google DeepMind's VLA that thinks in words before acting, with skills that transfer across different kinds of robots.

Google DeepMind released this on September 25, 2025, as an upgraded vision-language-action model (VLA) building on the original Gemini Robotics from March of that year. It adds two new capabilities. First, think-before-acting: before producing actions, it generates an internal chain of reasoning in natural language, breaking a semantically complex instruction down into simple steps. Second, Motion Transfer: a task trained only on the bimanual ALOHA 2 platform can run directly on the Apptronik Apollo humanoid or a dual-arm Franka setup, and vice versa — reusing skills across different robot embodiments. It works together with Gemini Robotics-ER 1.5, an embodied-reasoning model released at the same time: ER 1.5 handles high-level planning, directing 1.5 step by step in natural language. At launch, 1.5 was available only to select partners.

ExampleA task taught only on the ALOHA 2 bimanual platform can be carried out on the Apollo humanoid robot with no additional training.

Related
Gemini Robotics · Gemini Robotics-ER · Vision-Language-Action Model · Cross-Embodiment · Embodied Chain-of-Thought · Apptronik Apollo
Sources
Gemini Robotics 1.5 brings AI agents into the physical world (Google DeepMind)
As of
2025-09

See it in the full glossary →