Embodied AI Glossary中文

Visual Place Recognition

视觉位置识别VPRAdvanced

Looking at an image and determining which place it is, and whether the system has been there before.

Visual place recognition (VPR) takes a newly captured image and finds which place it was taken, by searching a pre-built database of images tagged with location — fundamentally an image retrieval problem. The difficulty is that the same place can look very different depending on lighting, season, time of day, and viewing angle. A common approach aggregates an image’s local features into a single global descriptor vector for similarity comparison; NetVLAD (CVPR 2016) uses a trainable aggregation layer, weakly supervised with Google Street View images from different years, and is a widely used baseline. In robotics, VPR mainly serves loop closure detection in SLAM and relocalization after tracking is lost (re-determining where the robot is on the map), and is also used in navigation based on topological maps.

ExampleA robot vacuum is picked up and set down in a different room; it compares the current view against the keyframes stored during mapping one by one, finds the closest match, and figures out which room it’s in.

Also called
VPR, Place Recognition
Related
Loop Closure Detection · Relocalization · Visual SLAM · Topological Map · Feature Matching · Navigation
Sources
Where is your place, Visual Place Recognition? (Garg et al., IJCAI 2021)
NetVLAD: CNN architecture for weakly supervised place recognition (CVPR 2016)

See it in the full glossary →