arXiv Computer Vision

AdaptVPR: Route-Aware Hard Positive Generation for Robust Visual Place Recognition

AdaptVPR introduces a route-aware generative augmentation framework that creates hard positive examples for Visual Place Recognition (VPR) training. It uses a vision‑language model to assess scene editability, a rule‑based scheduler to select generation routes, and a VPR‑oriented verification scheme to ensure geometric consistency and appearance diversity. The resulting AdaptCities dataset contains 160K verified synthetic hard positives, leading to consistent performance gains across VPR baselines, including up to 9.2% improvement in R@1 under challenging domain shifts.

arXiv Computer Vision
Sep 3

Domain shift-robust object detection with GenAI image editing

The paper investigates using diffusion-based generative image editing to improve object detector robustness against domain shifts, specifically camouflaged military vehicle detection. By synthetically adding foliage, netting, and multi‑spectral camouflage to training data with models such as Qwen Image Edit 2509 and Flux.2 Dev, the authors demonstrate significant mAP gains (up to +20.1 for foliage) over detectors trained on uncamouflaged data. LoRA fine‑tuning further boosts performance for the more challenging multi‑spectral camouflage.

By Isabel D. Stein, Thijs A. Eker, Sebastiaan P. Snel, Ella P. Fokkinga, Klamer Schutte, Luca Ambrogioni, Friso G. Heslinga
arXiv Machine Learning
Sep 22

MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments

arXiv:2505.12254v3 Announce Type: replace-cross Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...

By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv Computer Vision
Sep 21

Multi-viewpoint Geo-localization with Event Cameras

The paper presents MegaEvent, an event‑based visual place recognition system that remains robust to viewpoint changes. By converting five large‑scale geo‑tagged datasets into synthetic event streams and fine‑tuning a vision transformer with a multi‑loss function, MegaEvent achieves an average Recall@1 of 82% on three event‑based localization datasets, outperforming existing methods by 20 recall points. The authors also introduce the Springfield‑Event‑VPR dataset, a 3.7 km walking route recorded in three camera orientations, where MegaEvent surpasses the strongest baseline by 9 recall points.

By Adam D. Hines, Michael Milford, Tobias Fischer
arXiv AI
Sep 24

Geometry-Conditioned Visual Place Recognition in Natural Environments

The paper introduces Depth‑Aware Distillation (DAD), a method that conditions a pretrained Vision Foundation Model’s token representations on geometry inferred by a Geometric Foundation Model, without using a depth sensor. DAD projects image‑aligned depth into the VFM token space and selectively modulates visual representations through channel‑wise geometric conditioning. On the WildCross benchmark, DAD raises average inter‑sequence Recall@1 from 61.41% to 66.37% and Recall@5 from 65.86% to 72.49%, especially improving performance under reverse traversal and long‑term appearance variation.

By Walter Nedov, Saimunur Rahman, Kavindie Katuwandeniya, David Hall, Kaushik Roy, Peyman Moghadam