The paper introduces a pipeline that localizes frames from historical PTZ maritime video onto a reference panorama and then selects a context‑aware, diverse subset for ship detection training. By combining SuperPoint‑LightGlue localization, weather and solar‑state metadata, and diversity sampling, the method reduces 40,718 candidate frames to just 220 images for annotation. Fine‑tuned YOLOv26 on this compact set achieves high detection performance (AP50 ≈ 94.8%) while cutting annotation effort by 99.5%.
By Ignat Romanov, Andreas Hadjipieris, Neofytos Dimitriou
arXiv:2606. 30576v1 Announce Type: cross Abstract: Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.
By Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.
arXiv:2608.29852v1 Announce Type: new
Abstract: Automated maritime surveillance from satellite and aerial imagery requires large, precisely annotated datasets, which remain scarce for the instance-se...
By Amir Abbes, Ines Harrabi, Lucas Justin Yirepoa Kinda, Rim Trabelsi, Adnane Cabani, Fatma Abdelkefi
arXiv:2508. 04928v5 Announce Type: replace-cross Abstract: We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images.
By Rit Gangopadhyay, Jung-Hee Kim, Xien Chen, Patrick Rim, Hyoungseob Park, Alex Wong
AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.
By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu
Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most existing datasets predominantly focus on general o...
OpenCVL is a large, open dataset for fine-grained cross-view localization, comprising 617,388 ground‑aerial image pairs from 41 European cities. It blends high‑end sensor data with diverse in‑the‑wild images and includes a curation framework to correct pose annotations, enabling reliable evaluation. The dataset also offers cross‑area and snowy test sets to probe generalization, and experiments show that adding noisy in‑the‑wild data improves model performance on clean tests.
By Zimin Xia, Mubariz Zaffar, Junsheng Fu, Alexandre Alahi, Julian F. P. Kooij
arXiv:2608. 16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images.
By Zichao Zeng, Weijia Fan, Yufan Chen, June Moh Goo, Junwei Zheng, Ruiping Liu, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen, Jan Boehm
arXiv:2607.20116v2 Announce Type: replace
Abstract: Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, diff...
By Xin Li, Siyuan Duan, Shang Wang, Zhimin Mao, Bingliang Hu, Geng Zhang
PDA++ is a unified, environment‑aware object insertion framework for remote sensing imagery that improves few‑shot and long‑tail recognition. It operates in three stages: Planning, which selects scene‑compatible poses using an affordance field; Decoupling, which conditions the background on pose to preserve object identity while adapting to the scene; and Assimilation, which aligns multi‑scale texture distributions via optimal transport to enhance local coherence. The method achieves a whole‑image FID of 6.28 and boosts average few‑shot recognition mAP50 by 17.69 points on optical data, while also improving ship detection on SAR imagery and maintaining performance under cross‑dataset transfer and amorphous‑target insertion.
By Xianchi Dong, Yingyan Hou, Chao Ren, Wanxuan Lu, Zihan Wei, Hongfeng Yu, Yixiao Wang, Chubo Deng, Xian Sun
arXiv:2609.14462v1 Announce Type: new
Abstract: Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Exist...
By Jiaming Tan, Mingliang Zhai, Zhen Li, Yuwei Wu, Chuanhao Li, Kaipeng Zhang