Vision-and-Language Navigation
视觉语言导航VLNEssentialHaving an agent follow a natural-language route description and use vision to reach a destination in an unfamiliar space.
Vision-and-language navigation (VLN) is an embodied task proposed by Peter Anderson and colleagues in a CVPR 2018 paper: an agent gets an instruction like “exit and turn right, pass the sofa, and stop at the kitchen doorway,” and has to navigate step by step to the endpoint in a previously unvisited indoor environment, using only first-person vision. The companion Room-to-Room (R2R) dataset is built from real Matterport3D house scans, covering 90 buildings and about 22,000 human-written instructions. It tests how well language understanding, visual perception, and spatial memory work together. In the original R2R, the agent could only jump between waypoints with pre-rendered panoramas; a 2020 successor, VLN-CE, changed this to moving through continuous 3D space using low-level actions like moving forward and turning, closer to how a real robot operates. Models that run on actual robots, such as NaVid and NaVILA, followed after that.
ExampleGiven the instruction “walk straight down the hallway, turn left through the second door into the bedroom, and stop by the bed,” the robot looks and moves step by step, and succeeds if it stops by the bed.
- Also called
- VLN
- Related
- Navigation · Room-to-Room · Success weighted by Path Length · Object-Goal Navigation · NaVILA · Vision-Language Model
- Sources
- Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
Room-to-Room (R2R) 数据集主页 (Chinese)
Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments (VLN-CE)