EventGeM introduces a global‑to‑local feature fusion pipeline for event‑based visual place recognition, combining whole‑image feature detection with 2D homography‑based re‑ranking via RANSAC. It adds a regional generalized mean (GeM) pooling layer that learns to extract the most relevant spatial features from event streams, producing a compact global descriptor trained on the NYC‑Event‑VPR dataset. The method demonstrates significant improvements in viewpoint‑robust localization, achieving 7–43 percentage point gains in Recall@1 over the strongest baseline and real‑time performance on a robotic platform.
By Adam D. Hines, Gokul B. Nair, Nicol\'as Marticorena, Michael Milford, Tobias Fischer
arXiv:2608.23290v1 Announce Type: new
Abstract: Accurate visual localization on robotic and wearable platforms remains challenging in dense urban environments. Existing methodologies typically rely o...
By Antoni Valls, Jordi Sanchez-Riera
arXiv:2609.17168v1 Announce Type: cross
Abstract: Autonomous systems require reliable place recognition for efficient and effective simultaneous localisation and mapping (SLAM). Traditional geometric...
By Mayowa Adebambo, Sebastian Donnelly, Armand Amaritei, Andrew Bradley, Alexander Rast
The paper introduces Active Cross-View Object Geo-Localization (ActiveGeo), enabling mobile agents to actively select new viewpoints and decide when to stop to improve localization with fewer observations. It proposes the ActiveMoPT framework, which uses a three-stage training process: Multi-View Prompt-Preserving Adaptation, Trajectory-Guided Policy Initialization, and Cost-Aware Policy Refinement with GRPO. The authors also create a zero-shot test set, ActiveGeo-858, and demonstrate that ActiveMoPT outperforms prior methods on MoP-UAV and ActiveGeo-858.
By Shunyu Yao, Xiaohan Zhang, Zhuoran Yang, Haoqi Lai, Qi Ming, Xiaoxi Hu, Hui-Liang Shen, Si-Yuan Cao
VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.
By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang
arXiv:2606. 30576v1 Announce Type: cross Abstract: Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.
By Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu
arXiv:2608. 16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images.
By Zichao Zeng, Weijia Fan, Yufan Chen, June Moh Goo, Junwei Zheng, Ruiping Liu, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen, Jan Boehm
GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.
By Jinnao Li, Tingzhu Chen, Changbo Wang
The paper introduces CamVLM, a framework that equips large vision‑language models with the ability to actively control camera viewpoints for improved surveillance video understanding. It presents two new datasets: CCTV‑Anomaly, a large‑scale surveillance video collection with detailed captions and event annotations, and CamTrack‑53K, an object‑centric viewpoint trajectory dataset for learning camera actions. Using reinforcement learning, CamVLM learns long‑horizon observation strategies, achieving state‑of‑the‑art performance in both passive and dynamic viewpoint settings.
By Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Shichao Kan
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.
arXiv:2505.12254v3 Announce Type: replace-cross
Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...
By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv:2607.20116v2 Announce Type: replace
Abstract: Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, diff...
By Xin Li, Siyuan Duan, Shang Wang, Zhimin Mao, Bingliang Hu, Geng Zhang