EO-1
EO-1(EmbodiedOneVision)AdvancedA 3B open-source unified embodied model from Shanghai AI Lab where a single network both reasons in language and outputs actions.
EO-1 was released in August 2025, led by Shanghai AI Lab with real-robot support from AgiBot, with weights, training code, and data all open-source. It is built on Qwen2.5-VL-3B as a decoder-only Transformer: text is generated autoregressively, token by token, while actions are generated as continuous values by denoising through flow matching, with the two trained jointly inside the same model. Its companion dataset, EO-Data1.5M, contains 1.5 million interleaved “image-text-action” samples covering physical common sense, task reasoning, spatial understanding, and manipulation trajectories, letting the model reason and act within the same stretch of context. The company reports that it surpassed the then-current open-source models on ERQA, LIBERO, SimplerEnv, and its own EO-Bench.
ExampleReal-robot testing covered four robots — Franka, WidowX 250, AgiBot G-1, and the LeRobot SO-100 — on tasks including long-horizon dexterous manipulation and tasks that require reasoning before acting.
- Also called
- EmbodiedOneVision, EO1, EO-Robotics, An Open Unified Embodied Foundation Model for General Robot Control
- Related
- Vision-Language-Action Model · Hybrid Autoregressive-Diffusion Architecture · Flow Matching · Embodied Reasoning · Shanghai Artificial Intelligence Laboratory · LeRobot
- Sources
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control (arXiv 2508.21112)
EO-1 GitHub 仓库 (Chinese)
EO-1 项目页 (Chinese) - As of
- 2026-02