Mobility VLA
AdvancedA hierarchical navigation system that has a long-context VLM watch a tour video, then follows an image-and-text instruction to find the destination.
This was released by Google DeepMind in July 2024. The task it targets is called MINT (Multimodal Instruction Navigation with demonstration Tours): someone first walks through the environment with a camera to record a tour video, and afterward the user can give instructions with text plus an image — for example, holding an object and asking “where does this go back.” The system has two layers: the high level uses Gemini 1.5 Pro, with a context length of up to 1 million tokens, to read through the entire tour video and the instruction and locate the frame where the target is; the low level uses COLMAP (a tool that recovers camera pose from images) to build a topological map from the video — a map where locations are nodes and passable connections are edges — and generates waypoint actions from it for the base to execute. In a real, occupied 836-square-meter office, end-to-end success rates for instructions requiring reasoning and for multimodal instructions were 86% and 90%, respectively.
ExampleThe user holds up a charger and asks “where should this go back,” and the robot first locates, within the tour video, the desk where the charger belongs, then drives there along the topological map.
- Also called
- MINT (Multimodal Instruction Navigation with demonstration Tours), Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
- Related
- Vision-and-Language Navigation · Topological Map · Hierarchical Architecture · Google Gemini · Context Length · Structure from Motion
- Sources
- Mobility VLA (arXiv 2407.07775)
Mobility VLA (arXiv HTML full text) - As of
- 2024-07