arXiv:2609.37003v1 Announce Type: new
Abstract: Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most e...
By Danfeng Hong, Chenyu Li, Jocelyn Chanussot
arXiv:2608.29852v1 Announce Type: new
Abstract: Automated maritime surveillance from satellite and aerial imagery requires large, precisely annotated datasets, which remain scarce for the instance-se...
By Amir Abbes, Ines Harrabi, Lucas Justin Yirepoa Kinda, Rim Trabelsi, Adnane Cabani, Fatma Abdelkefi
Detecting vessels engaging in illegal activities is of paramount importance for maritime security. One of the major goals is to detect dark vessels, ships that disable their transponders to evade surveillance.
arXiv:2608. 09360v1 Announce Type: cross Abstract: The demand for maritime surveillance has given rise to the need for monitoring fishing vessel activities, particularly in addressing the challenge of "dark vessels" that operate without Automatic Identification System (AIS) transmission.
By Shantakar Mohanty, Prasun Kumar Gupta, Raian Vargas Maretto
The paper introduces a pipeline that localizes frames from historical PTZ maritime video onto a reference panorama and uses context-aware sampling to build compact, scene‑specific training sets for ship detection. By enriching frames with weather and solar metadata and applying diversity sampling, the method reduces 40,718 candidate frames to just 220 for annotation, achieving a 99.5% reduction. Fine‑tuned YOLOv26‑m on this curated subset attains high detection performance (AP50 ≈ 94.8% and AP50‑95 ≈ 75.1%).
The paper introduces a pipeline that localizes frames from historical PTZ maritime video onto a reference panorama and then selects a context‑aware, diverse subset for ship detection training. By combining SuperPoint‑LightGlue localization, weather and solar‑state metadata, and diversity sampling, the method reduces 40,718 candidate frames to just 220 images for annotation. Fine‑tuned YOLOv26 on this compact set achieves high detection performance (AP50 ≈ 94.8%) while cutting annotation effort by 99.5%.
By Ignat Romanov, Andreas Hadjipieris, Neofytos Dimitriou
arXiv:2607. 23537v1 Announce Type: new Abstract: Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality.
By Qiao Yan, Yihan Wang, Zhenghao Xing, Jiaqi Xu, Pheng-Ann Heng
arXiv:2608.21281v1 Announce Type: new
Abstract: Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging...
By Abigail G. Grassick, Jerome Tze-Hou Hsu, Ethan Lin, Ziang Liu, Max Whitton, Madelyn Hair, Liam Gutierrez, Haozheng Yu, Kristin Branson, Vivek Jayaraman, Michael A. Gil, Andrew M. Hein, Jennifer J. Sun
The paper introduces the Wide-area Spatio-temporal Scene Understanding (WSTU) problem, which demands simultaneous wide-area coverage, per-target resolution, and temporal continuity—capabilities lacking in existing datasets. To address this, the authors present HARD, an ultra‑high‑resolution (12768×9564) UAV dataset annotated for object detection, multi‑object tracking, and scene‑level visual question answering. They also propose a latency‑aware metric, streaming‑HOTA (s‑HOTA), and show through baseline experiments that high resolution and processing latency significantly impact detection, tracking, and VQA performance, revealing gaps in current methods for WSTU.
By Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang, Wenbin Li, Hongbing Pan
The paper introduces BMT, a unified hierarchical Vision Transformer that jointly performs SAR-to-optical image translation and semantic segmentation. It incorporates a LocalViTBlock, an enhanced output module, a ControlNet-style conditional injection, and a bounded Kendall uncertainty weighting scheme to balance the two tasks. Experiments on paired and unpaired datasets demonstrate competitive performance in both translation quality and segmentation accuracy.
By Siyuan Liu, Xuze Zhang, Yongshun Wang, Licong Pan, Hang Liu, Huihui Li
KSG‑Net introduces a Key‑Sparse and Global‑Context learning framework for maritime 3D ship detection, addressing weak feature representation of small, sparse vessels and limited global modeling of large vessels. The network employs a Key Sparse Multi‑scale Aggregation module to select informative voxels and aggregate cross‑scale features, and a Global Context Aggregation module to capture long‑range geometric dependencies via scene‑level context modeling. Experiments on the Thames River vessel dataset and simulated data show that KSG‑Net outperforms existing methods in multi‑scale vessel detection and remains robust in complex maritime environments.
By Zhouyuan Huai, Meiqi Wan, Yan Yang, Minshi Chen, Xin Yuan, Wei Wang, Xiao Wang
arXiv:2608.24594v1 Announce Type: new
Abstract: Submerged kelp forests are vital coastal ecosystems that support marine biodiversity and ecosystem dynamics, yet accurate underwater kelp segmentation...
By Sundarabalan Balasubramanian, C\'esar Borja, Ana C. Murillo, Lexi N. Wilkes, Meredith L. McPherson, Kira A. Krumhansl, Jennifer A. Dijkstra, Jarrett E. K. Byrnes