Embodied AI Glossary中文

AgiBot GO-1

智元 GO-1(启元大模型)GO-1Common

AgiBot's 2025 general-purpose embodied foundation model, using latent actions so human video can help train it too.

GO-1 is the general-purpose manipulation policy AgiBot released on March 10, 2025, alongside the technical report for its AgiBot World dataset. It proposes the ViLLA (vision-language-latent-action) framework in three parts: a latent action model compresses the change between adjacent frames into discrete latent action tokens, and can be trained on human videos with no action labels, such as Ego4D; a latent planner, built on the InternVL2.5-2B vision-language model, predicts these tokens from images and instructions; an action expert then uses a diffusion objective to output continuous, high-frequency actions. The paper reports that pretraining on AgiBot World — over 1 million trajectories collected from 100 robots — improves average performance 30% over using Open X-Embodiment; on complex real-robot tasks GO-1 reaches over 60% success, 32 points higher than RDT. The model was open-sourced in September 2025, together with a lighter GO-1 Air variant that drops the latent planner; its successor is 2026's GO-2.

ExampleThe open-sourced GO-1 weights are hosted on Hugging Face; the official documentation says inference needs about 7GB of GPU memory, and full-parameter fine-tuning with batch size 16 needs about 70GB.

Also called
Genie Operator-1, GO-1 Air
Related
Vision-Language-Latent-Action · Latent Action · AgiBot World · AgiBot GO-2 · Action Expert · AgiBot
Sources
AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems (arXiv 2503.06669)
OpenDriveLab/AgiBot-World(GitHub,含 GO-1 / GO-1 Air 开源说明) (Chinese)
As of
2025-09

See it in the full glossary →