LM-Nav
AdvancedCombines GPT-3, CLIP, and a visual navigation model so a robot can navigate by following natural-language instructions.
LM-Nav was released in July 2022 by Dhruv Shah, Brian Ichter, Sergey Levine, and colleagues, published at CoRL 2022. Directing a robot's navigation with language usually needs a lot of trajectory data with text descriptions, which is expensive to annotate. LM-Nav does no fine-tuning at all and uses no language-labeled robot data; instead, it combines three off-the-shelf pretrained models: the large language model GPT-3 breaks the instruction down into a sequence of landmarks; the image-text model CLIP judges which landmark corresponds to what the robot's camera sees; and the visual navigation model ViNG builds a topological map of the environment from previously collected images and runs a go-to-point policy. The system then searches for the shortest route that passes through those landmarks in order. It completed long-distance navigation in real outdoor environments, and is an early representative example of assembling a robot system out of foundation models.
ExampleFor example, the user says “go past the stop sign and head to the white building”; GPT-3 extracts the two landmarks “stop sign” and “white building,” CLIP locates the corresponding positions in the topological map, and ViNG drives to them in sequence.
- Also called
- Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
- Related
- Vision-and-Language Navigation · CLIP · Topological Map · LLM-based Task Planning · GNM · ViNT
- Sources
- arXiv 2207.04429: LM-Nav
LM-Nav 项目主页 (Chinese) - As of
- 2022-07