The paper introduces a pipeline that localizes frames from historical PTZ maritime video onto a reference panorama and uses context-aware sampling to build compact, scene‑specific training sets for ship detection. By enriching frames with weather and solar metadata and applying diversity sampling, the method reduces 40,718 candidate frames to just 220 for annotation, achieving a 99.5% reduction. Fine‑tuned YOLOv26‑m on this curated subset attains high detection performance (AP50 ≈ 94.8% and AP50‑95 ≈ 75.1%).
arXiv:2606. 30576v1 Announce Type: cross Abstract: Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.
By Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu
AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.
By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu
arXiv:2609.14462v1 Announce Type: new
Abstract: Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Exist...
By Jiaming Tan, Mingliang Zhai, Zhen Li, Yuwei Wu, Chuanhao Li, Kaipeng Zhang
arXiv:2608.29852v1 Announce Type: new
Abstract: Automated maritime surveillance from satellite and aerial imagery requires large, precisely annotated datasets, which remain scarce for the instance-se...
By Amir Abbes, Ines Harrabi, Lucas Justin Yirepoa Kinda, Rim Trabelsi, Adnane Cabani, Fatma Abdelkefi
arXiv:2508. 04928v5 Announce Type: replace-cross Abstract: We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images.
By Rit Gangopadhyay, Jung-Hee Kim, Xien Chen, Patrick Rim, Hyoungseob Park, Alex Wong