arXiv:2608.28891v1 Announce Type: new
Abstract: Pixel-level cross-view geo-registration aims to align a query image (e.g., drone) to a geo-referenced satellite map so that every query pixel can be ma...
By Qingyang Liu, David G Shatwell, Parth Parag Kulkarni, Mubarak Shah
UniGeo is a multimodal large language model designed for text-guided drone geo‑localization, enabling the identification of target regions in large image galleries from natural‑language descriptions. It integrates geo‑semantic understanding, cross‑view semantic generation, and candidate‑level verification within a shared vision‑language framework, establishing stable correspondences among local scene elements, spatial relations, and language. A multi‑stage training strategy progressively refines geo‑semantic learning, cross‑view mapping, and fine‑grained verification, yielding significant performance gains on GeoText‑1652, with R@10 and mAP improvements of 13.59 and 2.83 percentage points respectively.
By Jiahao Wen, Hang Yu, Zhedong Zheng
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.
arXiv:2604.13183v4 Announce Type: replace
Abstract: Generalizable cross-view geo-localization aims to match the same location across views in unseen regions and conditions without GPS supervision. It...
By Hongyang Zhang, Yinhao Liu, Haitao Zhang, Zhongyi Wen, Zhenyu Kuang, Shuxian Liang, Xian-Sheng Hua
arXiv:2607.20116v2 Announce Type: replace
Abstract: Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, diff...
By Xin Li, Siyuan Duan, Shang Wang, Zhimin Mao, Bingliang Hu, Geng Zhang
arXiv:2608.29211v1 Announce Type: new
Abstract: Ground image localization with respect to satellite imagery is a key enabler for metrically-accurate, geo-localized 3D scene reconstruction from uncons...
By Angel Daruna, Ben Southall, Niluthpol Chowdhury Mithun, Kshitij Minhas, Nicholas Meegan, Qiao Wang, Bogdan Matei, Supun Samarasekera, Rakesh Kumar
AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.
By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu
arXiv:2512.02697v4 Announce Type: replace
Abstract: Cross-view geo-localization infers a location by retrieving geo-tagged reference images matching a query image. However, the traditional satellite-...
By Zixuan Song, Jing Zhang, Di Wang, Zhiming Luo, Wenbin Liu, Haonan Guo, En Wang, Bo Du, Liangpei Zhang
M3GA-Wild is a new benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests, combining synchronized RGB imagery and LiDAR from ground traversals with high‑resolution aerial imagery and multi‑altitude LiDAR over 370 hectares. The dataset includes accurate geo‑referenced 6‑DoF poses and spans 36 km of forest traversals, enabling systematic evaluation of visual, LiDAR, cross‑modal, and multi‑modal methods. Baseline experiments show LiDAR outperforms vision‑only approaches under severe viewpoint changes, while current multi‑modal fusion offers limited gains due to poor cross‑modal alignment, highlighting challenges in cross‑platform localisation and domain gaps.
By Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani
arXiv:2608. 09270v1 Announce Type: cross Abstract: Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation.
By Jiahui Cui, Yan Zhao, Kan Wei, Enze Zhu, Peirong Zhang, Lei Wang, Yiru Wang
arXiv:2606. 07708v1 Announce Type: cross Abstract: We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections.
By Prakhar Bhardwaj, Simone Weikl, Kilian Mang, Elia Jonas Sandtner
arXiv:2609.13903v1 Announce Type: new
Abstract: We study how to update a pre-built aerial scene with a newly captured, unposed ground-view sequence. The aerial scene already contains a reliable metri...
By Zhuoxiao Li, Xinyi Liu, Taoyu Wu, Yinrui Ren, Tongyan Hua, Ou Jing, Shuai Zhang, Dongli Wu, Rongjun Qin, Ge Lin Kan, Wufan Zhao