Embodied AI Glossary中文

NavFoM (Galbot)

银河通用 NavFoMNavFoMAdvanced

A 2025 navigation foundation model from Galbot and Peking University, with one set of weights adapting to many embodiments and navigation tasks.

NavFoM was proposed by Galbot together with Peking University, the University of Adelaide, Zhejiang University, and others; the paper went up on arXiv in September 2025, and the company formally launched it in November, promoted as “the world's first cross-embodiment, all-around navigation foundation model.” Earlier navigation models were mostly trained separately, one robot and one task at a time; NavFoM instead uses 8.02 million navigation samples (quadruped, drone, wheeled robot, car) plus 4.76 million image-text and video question-answer samples to simultaneously learn instruction navigation, object finding, target tracking, and autonomous driving. It uses Qwen2-7B as its language backbone, concatenating DINOv2 and SigLIP visual features, with special tokens marking which camera and which moment each frame comes from, supporting 1 to 8 camera feeds; an MLP finally outputs waypoints, handed off to the embodiment's local planner for execution.

ExampleThe same weights, with no task-specific fine-tuning, are evaluated on benchmarks including VLN-CE instruction navigation, HM3D-OVON open-vocabulary object finding, EVT-Bench target tracking, and NAVSIM autonomous driving; Peking University and Galbot's later UrbanVLA was also trained on top of it.

Also called
Embodied Navigation Foundation Model
Related
Cross-Embodiment · Vision-and-Language Navigation · Embodied Visual Tracking · TrackVLA · NaVid · AstraBrain
Sources
Embodied Navigation Foundation Model (arXiv 2509.12129)
NavFoM 项目页 (Chinese)
银河通用发布全球首个跨本体全域环视导航大模型NavFoM(网易,2025-11-05) (Chinese)
As of
2025-11

See it in the full glossary →