Embodied AI Glossary中文

InternVL

书生 InternVLAdvanced

Shanghai AI Lab's open-source vision-language model series, starting from a 6-billion-parameter vision encoder and scaling up from there.

InternVL is an open-source vision-language model series from Shanghai AI Lab's OpenGVLab team, with the Chinese name “Shusheng Wanxiang.” The original paper was released in December 2023 and selected as an oral presentation at CVPR 2024; the idea was to scale up the vision encoder, InternViT, to 6 billion parameters, then progressively align it with a large language model, evaluated across 32 vision-language benchmarks. Later versions — 1.5, 2, 2.5, 3, and 3.5 — kept the same “vision encoder + MLP projection layer + large language model” structure, ranging in size from 1 billion to 241 billion total parameters, with code open-sourced under the MIT license; as of September 2026, the 3.5 series is the latest. In embodied AI, it's commonly used as the vision-language backbone for VLA (vision-language-action) models.

ExampleAgiBot GO-1's latent planner uses InternVL2.5-2B as its backbone: it reads in multiple camera feeds and an instruction, first predicts latent action tokens, then hands them to an action expert to generate continuous actions.

Also called
OpenGVLab InternVL, InternVL-Chat
Related
Vision-Language Model · Multimodal Large Language Model · Shanghai Artificial Intelligence Laboratory · InternVL · InternVLA (Shanghai AI Laboratory) · AgiBot GO-1
Sources
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks (arXiv 2312.14238)
OpenGVLab/InternVL (GitHub)
As of
2026-09

See it in the full glossary →