NaVILA
AdvancedA legged-robot navigation VLA where a large model gives mid-level actions in text and an RL locomotion controller does the walking.
NaVILA was proposed by Xiaolong Wang's group at UC San Diego together with NVIDIA and USC, posted to arXiv in December 2024, published at RSS 2025. A vision-language model is good at understanding images and instructions, but outputting leg joint commands directly is hard, so NaVILA splits into two layers: the high level is a VLA fine-tuned from NVIDIA's VILA (8B), which looks at the camera view and the instruction and outputs, in text, a mid-level action with distance and angle, such as “move forward 75 centimeters”; the low level is a locomotion policy trained with reinforcement learning that reads a lidar-generated height map and turns that mid-level action into commands for a quadruped's 12 joints. Besides simulated navigation data, training data also includes about 2,000 first-person YouTube tour videos. It reaches a 54% success rate on R2R-CE, and the paper also released VLN-CE-Isaac, a benchmark built on Isaac Lab.
ExampleIn real-robot tests across 25 instructions spanning office, home, and outdoor settings, NaVILA running on a Unitree Go2 quadruped reached 88% success, and 75% on complex multi-room instructions; the same model also works on a Booster T1 humanoid with no retraining.
- Also called
- Legged Robot Vision-Language-Action Model for Navigation
- Related
- Vision-and-Language Navigation · Vision-Language-Action Model · Legged Locomotion · Hierarchical Architecture · RL-based Locomotion Control · NaVid
- Sources
- NaVILA: Legged Robot Vision-Language-Action Model for Navigation (arXiv 2412.04453)
NaVILA 项目页 (Chinese) - As of
- 2025-06