NoMaD
AdvancedA 2023 Berkeley navigation diffusion policy where one model can both explore freely and head toward a goal image.
NoMaD was proposed in October 2023 by Sergey Levine's group at UC Berkeley (Ajay Sridhar, Dhruv Shah, and others), published at ICRA 2024, where the project page says it won that year's best paper award. Robot navigation commonly needs two things: goal-free exploration in an unfamiliar environment, and heading toward a goal once a photo of it is seen. Previously this took two separate models; NoMaD merges them with “goal masking”: half of training samples have the goal image masked out, and at inference, masking it produces exploration while revealing it produces goal-reaching. It uses a ViNT-style Transformer to encode recent frames and the goal image, then a diffusion model to generate a future sequence of waypoints, able to express multiple viable options at a fork in the path (action multimodality). The model has about 19 million parameters, and was trained on more than 100 hours of real, multi-robot data from GNM, SACSoN, and others.
ExampleExploring real environments on a LoCoBot, NoMaD reaches a 98% success rate with an average of 0.2 collisions; a compared baseline that first generates a sub-goal image with diffusion and then navigates gets 77% and 1.7 collisions, while also having about 15 times more parameters.
- Also called
- Goal Masked Diffusion Policies for Navigation and Exploration
- Related
- Diffusion Policy · ViNT · GNM · Image-Goal Navigation · Active Exploration · Action Multimodality
- Sources
- NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration (arXiv 2310.07896)
NoMaD 项目页 (Chinese) - As of
- 2024-05