ViNT
AdvancedA visual navigation foundation model from Berkeley, trained on navigation data from many kinds of robots, that finds a goal from an image.
ViNT was released by Dhruv Shah, Sergey Levine, and colleagues at UC Berkeley in June 2023, an oral-presentation paper at CoRL 2023. Earlier navigation models were mostly trained on data from a single robot in a single kind of environment. ViNT encodes images with EfficientNet followed by a Transformer, taking in the current frame plus several recent past frames along with a goal image, and predicting how far away the goal is and what to do next; its training data comes from multiple robot platforms, totaling hundreds of hours of navigation data. Paired with a diffusion model that generates candidate subgoal images, it can explore unfamiliar environments, and combined with long-range heuristics like GPS it can perform kilometer-scale navigation; it can also be adapted through prompt tuning to take GPS waypoints or turn-by-turn instructions as the goal instead. ViNT builds on the earlier GNM, and was itself followed by NoMaD.
ExampleA robot first drives along a route taking a series of photos; later, given just one of those photos as a goal image, ViNT can navigate back to the spot where that photo was taken in the same environment, and the same model can be deployed on different mobile robots.
- Also called
- Visual Navigation Transformer, ViNT: A Foundation Model for Visual Navigation
- Related
- GNM · NoMaD · Image-Goal Navigation · Navigation · Foundation Model · Prompt Tuning / Soft Prompt
- Sources
- ViNT (arXiv 2306.14846)
ViNT project page - As of
- 2023-10