arXiv:2606. 02341v1 Announce Type: cross Abstract: Underwater acoustic classification has a wide array of oceanic applications, but faces challenges due to an increasingly complex acoustic environment.
By Amirmohammad Mohammadi, Joshua Peeples, Alexandra Van Dine
SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.
By Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
arXiv:2608. 19710v1 Announce Type: cross Abstract: Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur.
By Mohammad Arif Ul Alam
The paper presents a three‑stage, parameter‑efficient approach to improve automatic target recognition (ATR) with synthetic aperture sonar (SAS) data by adapting DINOv3 Vision Transformers. Stage 1 applies Low‑Rank Adaptation (LoRA) while freezing the backbone, which significantly boosts the area under the precision‑recall curve from 0.300 to 0.679. Subsequent hard‑negative mining and supervised contrastive learning stages show negligible impact, indicating that a single LoRA adaptation is sufficient for effective underwater ATR.
By Dan Zimmerman, Frank E. Bobe III, Amelia L. McCormack, Matthew Cook, Gregory D. Vetaw
The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.
By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.
By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
arXiv:2601.08358v2 Announce Type: replace
Abstract: Even though the ocean covers the majority of the planet's surface, it remains the least explored ecosystem. As light and radio waves do not propaga...
By Hilde I. Hummel, Sandjai Bhulai, Rob D. van der Mei, Burooj Ghani
The paper addresses the spectral bias of Fourier Neural Operators (FNO) in predicting underwater acoustic transmission loss. It introduces a Spectral‑Spatial Residual Learning (S2RL) framework that first uses a spectral global propagator for coarse predictions and then a spatial local refiner to recover high‑frequency details. Experiments on a South China Sea dataset show that S2RL outperforms FNO baselines while keeping inference times in the millisecond range.
By Yifan Sun, Shikai Fang, Chao Zhang, Lei Cheng, Jianlong Li, Peter Gerstoft
RSFusionDet introduces a new RGB‑Sonar multimodal object detection dataset (RSFusion) and benchmark metrics for underwater imaging. The proposed detector uses a Cross‑Attention Fusion module to align RGB and sonar features and an Object Matching Head with loss to identify identical objects across modalities. On the RSFusion dataset, RSFusionDet achieves 76.4/48.6 AP for RGB/sonar detection and 83.4 F1‑Score for cross‑modal matching, outperforming existing models and improving over the DINO baseline by 0.7/1.4 AP.
By Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang, Ye Li
arXiv:2607. 08031v1 Announce Type: cross Abstract: The dynamics of communication environments induce significant distribution shifts across domains, challenging the generalization of deep learning-based automatic modulation classification (AMC) models.
By Shuang Wang, Chenxu Wang, Hantong Xing, Hanlin Mo, Lirong Han, Licheng Jiao
arXiv:2606. 05754v1 Announce Type: cross Abstract: Phase-sensitive optical time-domain reflectometry ($\phi$-OTDR) is widely used in large-scale distributed acoustic sensing (DAS) because it provides distributed spatiotemporal monitoring over long sensing distances.
By Weiguang Wang, Fugen Wu, Hailing Wang, Xuechen Liang, Xiaobin Li, Ru Han, Tianchang Xie
arXiv:2509. 04682v2 Announce Type: replace-cross Abstract: Deploying reliable bioacoustic monitoring systems requires models that generalize under high-noise, low-SNR conditions and evaluation protocols that expose deployment-relevant failure modes, gaps largely unaddressed in current UPAM practice.
By Nicholas R. Rasmussen, Rodrigue Rizk, Longwei Wang, KC Santosh