Embodied AI Glossary中文

Aerial Vision-and-Language Navigation

空中视觉语言导航(无人机 VLN)Aerial VLNAdvanced

Having a drone understand a natural-language instruction and fly to a target location in a 3D outdoor space such as a city.

Vision-and-language navigation (VLN) originally studied ground robots walking indoors by following language instructions; aerial VLN brings this to drones: the agent sees a first-person view and flies outdoors following an instruction describing landmarks and a route. Compared with the ground setting, it adds an extra dimension, altitude; paths routinely run hundreds of meters, instructions have to reference more landmarks, and both localization and long-range memory become harder. A representative benchmark is AerialVLN (ICCV 2023), built on Unreal Engine 4 and Microsoft AirSim across 25 city-scale scenes, collecting 8,446 flight paths and 25,338 instructions with an average path length of 661.8 meters; actions include moving forward, turning left or right, ascending, descending, strafing left or right, and stopping.

ExampleOn the AerialVLN test set, following instructions averaging 83 English words, the CMA baseline reaches a success rate of only 1.6%, compared with 80.8% for humans — a very large gap.

Also called
UAV VLN, Aerial VLN
Related
Vision-and-Language Navigation · Unmanned Aerial Vehicle (UAV) · Room-to-Room · Navigation · Long-horizon Task · Aerial Manipulation
Sources
AerialVLN: Vision-and-Language Navigation for UAVs (arXiv 2308.06735, ICCV 2023)

See it in the full glossary →