Embodied AI Glossary中文

NVIDIA Eagle VLM

Eagle(英伟达 VLM)Advanced

NVIDIA's open-source vision-language model series, used as the vision-language backbone in GR00T N1 through N1.6.

Eagle is an open-source vision-language model (VLM) series from NVIDIA's research team. The original Eagle (August 2024) studied how to combine multiple vision encoders, finding that simply concatenating the visual tokens from several complementary encoders worked well; Eagle 2 (January 2025) published the details of its post-training data strategy, with sizes including 1B, 2B, and 9B; Eagle 2.5 (technical report released April 2025) targets long context, supporting up to 128K tokens and strengthening long-video and high-resolution image understanding. Its main role in embodied AI is as the 'System 2' vision-language backbone for GR00T: GR00T N1 uses Eagle 2, while N1.5 and N1.6 use improved versions from the Eagle 2.5 family. According to the official GR00T repository, N1.7 onward switches to Cosmos-Reason2-2B instead.

ExampleGR00T N1.5 starts from Eagle 2.5 and fine-tunes it further for object localization and physical understanding, while keeping this VLM frozen during both pretraining and fine-tuning so its parameters never update.

Also called
Eagle 2, Eagle 2.5, NVEagle
Related
NVIDIA Isaac GR00T N1 · Vision-Language Model · Vision Encoder · Dual-System Architecture (System 1 / System 2) · NVIDIA Cosmos Reason · Vision-Language-Action Model
Sources
GitHub: NVlabs/EAGLE
NVIDIA GEAR: GR00T N1.5
GitHub: NVIDIA/Isaac-GR00T
As of
2026-09

See it in the full glossary →