arXiv AI

YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition

YILDIZ-VPR is a new visual geo‑localization dataset collected by repeatedly walking across the Davutpasa campus of Yildiz Technical University. It offers dense coverage of outdoor scenes captured at various times of day, seasons, and weather conditions, featuring historical buildings, modern structures, roads, green areas, and wooded regions. Each GoPro 9 video is synchronized with GPS, gyroscope, speed, and temperature data, providing precise location labels for extracted frames.

arXiv Machine Learning
Sep 22

MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments

arXiv:2505.12254v3 Announce Type: replace-cross Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...

By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv Computer Vision
Sep 21

Multi-viewpoint Geo-localization with Event Cameras

The paper presents MegaEvent, an event‑based visual place recognition system that remains robust to viewpoint changes. By converting five large‑scale geo‑tagged datasets into synthetic event streams and fine‑tuning a vision transformer with a multi‑loss function, MegaEvent achieves an average Recall@1 of 82% on three event‑based localization datasets, outperforming existing methods by 20 recall points. The authors also introduce the Springfield‑Event‑VPR dataset, a 3.7 km walking route recorded in three camera orientations, where MegaEvent surpasses the strongest baseline by 9 recall points.

By Adam D. Hines, Michael Milford, Tobias Fischer
arXiv Computer Vision
Aug 27

OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization

OpenCVL is a large, open dataset for fine-grained cross-view localization, comprising 617,388 ground‑aerial image pairs from 41 European cities. It blends high‑end sensor data with diverse in‑the‑wild images and includes a curation framework to correct pose annotations, enabling reliable evaluation. The dataset also offers cross‑area and snowy test sets to probe generalization, and experiments show that adding noisy in‑the‑wild data improves model performance on clean tests.

By Zimin Xia, Mubariz Zaffar, Junsheng Fu, Alexandre Alahi, Julian F. P. Kooij
arXiv Computer Vision
Aug 27

GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction

GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.

By Jinnao Li, Tingzhu Chen, Changbo Wang
arXiv Machine Learning
Aug 21

From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.

By Joan Perez, Giovanni Fusco
arXiv Computer Vision
Sep 21

EventGeM: Global-to-Local Feature Matching for Event-Based Visual Place Recognition

EventGeM introduces a global‑to‑local feature fusion pipeline for event‑based visual place recognition, combining whole‑image feature detection with 2D homography‑based re‑ranking via RANSAC. It adds a regional generalized mean (GeM) pooling layer that learns to extract the most relevant spatial features from event streams, producing a compact global descriptor trained on the NYC‑Event‑VPR dataset. The method demonstrates significant improvements in viewpoint‑robust localization, achieving 7–43 percentage point gains in Recall@1 over the strongest baseline and real‑time performance on a robotic platform.

By Adam D. Hines, Gokul B. Nair, Nicol\'as Marticorena, Michael Milford, Tobias Fischer