Embodied AI Glossary中文

Embodied-R1

Advanced

A 3B embodied-reasoning model from Tianjin University that uses “pointing” as an intermediate representation, trained with reinforcement fine-tuning.

Embodied-R1 is work released in August 2025 by a team at Tianjin University, accepted to ICLR 2026. The authors call the gap between “understanding what to do” and “actually doing it” the seeing-to-doing gap, and propose using “pointing” — outputting a point, a region, or a sequence of trajectory points on an image — as an intermediate representation that is independent of any specific robot body, defining four abilities: referring-expression grounding, region grounding, functional-part grounding, and visual-trajectory generation. The model is built on Qwen2.5-VL-3B, and is trained with two-stage reinforcement fine-tuning on a self-built dataset, Embodied-Points-200K, using the GRPO algorithm with automatically scored, task-specific rewards. The points and trajectories the model outputs are then handed off to lower-level modules such as motion planning for execution. The weights and dataset are both open-source.

ExampleWithout any task-specific fine-tuning, Embodied-R1 reaches a 56.2% success rate in SimplerEnv simulation and 87.5% across 8 real-robot xArm tasks, which the paper reports as a 62% improvement over strong baselines.

Also called
Embodied R1, Reinforced Embodied Reasoning for General Robotic Manipulation
Related
Embodied Reasoning Model · Pointing · Reinforcement Fine-Tuning (RL Fine-Tuning) · Group Relative Policy Optimization · Intermediate Representation · Affordance
Sources
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation (arXiv 2508.13998)
Embodied-R1 GitHub 仓库 (Chinese)
As of
2026-03

See it in the full glossary →