Embodied AI Glossary中文

Pelican-VL

北京人形 Pelican-VLAdvanced

An open-source vision-language 'brain' model for embodied robots, released by China's X-Humanoid center.

Pelican-VL is an embodied 'brain' model from the Beijing Humanoid Robot Innovation Center (X-Humanoid). The technical report came out in late October 2025, with weights open-sourced in November 2025; it is built on Qwen2.5-VL, and the first release shipped 7B and 72B parameter versions under the Apache 2.0 license. Rather than outputting joint-level actions directly, Pelican-VL handles spatial understanding, affordance reasoning, task planning, and function calling, and issues instructions to a lower-level controller or a vision-language-action (VLA) model. Its training method, DPPO (Deliberate Practice Policy Optimization), first uses GRPO (Group Relative Policy Optimization) reinforcement learning to find cases the model handles poorly, turns those hard cases into supervised fine-tuning data, and alternates between RL and SFT in a loop. The report claims a 20.3% average improvement over the base model, and leads open-source models of similar size on embodied benchmarks such as Where2Place and RefSpatialBench. Its Hugging Face page later added 3B and 235B-A22B versions as well.

ExampleIn one real-robot experiment from the report, the model predicts and continuously corrects grip force from the camera feed to complete a contact-rich grasp; in another, it acts as a unified 'brain' coordinating several different robots on a long-horizon task.

Also called
Pelican-VL 1.0, Pelican1.0-VL, Pelican-VL: A Foundation Brain Model for Embodied Intelligence
Related
Embodied Reasoning Model · Vision-Language Model · Qwen-VL · Group Relative Policy Optimization · Brain–Cerebellum Architecture · Beijing Humanoid Robot Innovation Center
Sources
arXiv 2511.00108: Pelican-VL 1.0
Hugging Face: X-Humanoid/Pelican1.0-VL-7B
Hugging Face: X-Humanoid 组织页 (Chinese)
As of
2026-09

See it in the full glossary →