arXiv Computer Vision

Multi-viewpoint Geo-localization with Event Cameras

The paper presents MegaEvent, an event‑based visual place recognition system that remains robust to viewpoint changes. By converting five large‑scale geo‑tagged datasets into synthetic event streams and fine‑tuning a vision transformer with a multi‑loss function, MegaEvent achieves an average Recall@1 of 82% on three event‑based localization datasets, outperforming existing methods by 20 recall points. The authors also introduce the Springfield‑Event‑VPR dataset, a 3.7 km walking route recorded in three camera orientations, where MegaEvent surpasses the strongest baseline by 9 recall points.

arXiv Computer Vision
Sep 21

EventGeM: Global-to-Local Feature Matching for Event-Based Visual Place Recognition

EventGeM introduces a global‑to‑local feature fusion pipeline for event‑based visual place recognition, combining whole‑image feature detection with 2D homography‑based re‑ranking via RANSAC. It adds a regional generalized mean (GeM) pooling layer that learns to extract the most relevant spatial features from event streams, producing a compact global descriptor trained on the NYC‑Event‑VPR dataset. The method demonstrates significant improvements in viewpoint‑robust localization, achieving 7–43 percentage point gains in Recall@1 over the strongest baseline and real‑time performance on a robotic platform.

By Adam D. Hines, Gokul B. Nair, Nicol\'as Marticorena, Michael Milford, Tobias Fischer
arXiv Computer Vision
Sep 18

Towards Active Cross-View Object Geo-Localization

The paper introduces Active Cross-View Object Geo-Localization (ActiveGeo), enabling mobile agents to actively select new viewpoints and decide when to stop to improve localization with fewer observations. It proposes the ActiveMoPT framework, which uses a three-stage training process: Multi-View Prompt-Preserving Adaptation, Trajectory-Guided Policy Initialization, and Cost-Aware Policy Refinement with GRPO. The authors also create a zero-shot test set, ActiveGeo-858, and demonstrate that ActiveMoPT outperforms prior methods on MoP-UAV and ActiveGeo-858.

By Shunyu Yao, Xiaohan Zhang, Zhuoran Yang, Haoqi Lai, Qi Ming, Xiaoxi Hu, Hui-Liang Shen, Si-Yuan Cao
arXiv Computer Vision
Sep 21

VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.

By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang
arXiv Computer Vision
Aug 27

GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction

GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.

By Jinnao Li, Tingzhu Chen, Changbo Wang
arXiv Computer Vision
3d ago

Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding

The paper introduces CamVLM, a framework that equips large vision‑language models with the ability to actively control camera viewpoints for improved surveillance video understanding. It presents two new datasets: CCTV‑Anomaly, a large‑scale surveillance video collection with detailed captions and event annotations, and CamTrack‑53K, an object‑centric viewpoint trajectory dataset for learning camera actions. Using reinforcement learning, CamVLM learns long‑horizon observation strategies, achieving state‑of‑the‑art performance in both passive and dynamic viewpoint settings.

By Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Shichao Kan
Hugging Face Trending Papers
Jul 22

RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.

arXiv Machine Learning
Sep 22

MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments

arXiv:2505.12254v3 Announce Type: replace-cross Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...

By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini