Embodied AI Glossary中文

ViNT

Advanced

A visual navigation foundation model from Berkeley, trained on navigation data from many kinds of robots, that finds a goal from an image.

ViNT was released by Dhruv Shah, Sergey Levine, and colleagues at UC Berkeley in June 2023, an oral-presentation paper at CoRL 2023. Earlier navigation models were mostly trained on data from a single robot in a single kind of environment. ViNT encodes images with EfficientNet followed by a Transformer, taking in the current frame plus several recent past frames along with a goal image, and predicting how far away the goal is and what to do next; its training data comes from multiple robot platforms, totaling hundreds of hours of navigation data. Paired with a diffusion model that generates candidate subgoal images, it can explore unfamiliar environments, and combined with long-range heuristics like GPS it can perform kilometer-scale navigation; it can also be adapted through prompt tuning to take GPS waypoints or turn-by-turn instructions as the goal instead. ViNT builds on the earlier GNM, and was itself followed by NoMaD.

ExampleA robot first drives along a route taking a series of photos; later, given just one of those photos as a goal image, ViNT can navigate back to the spot where that photo was taken in the same environment, and the same model can be deployed on different mobile robots.

Also called
Visual Navigation Transformer, ViNT: A Foundation Model for Visual Navigation
Related
GNM · NoMaD · Image-Goal Navigation · Navigation · Foundation Model · Prompt Tuning / Soft Prompt
Sources
ViNT (arXiv 2306.14846)
ViNT project page
As of
2023-10

See it in the full glossary →