arXiv AI

Heterogeneous 2D/1D Signal Representation Fusion for Underwater Acoustic Modulation Recognition Under Distribution Shift

arXiv:2606. 23702v1 Announce Type: cross Abstract: Modulation recognition systems rely on heterogeneous signal representations.

arXiv AI
Aug 26

SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.

By Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
arXiv AI
Sep 21

LoRA Enhanced Contrastive Learning with SAS Vision Transformers

The paper presents a three‑stage, parameter‑efficient approach to improve automatic target recognition (ATR) with synthetic aperture sonar (SAS) data by adapting DINOv3 Vision Transformers. Stage 1 applies Low‑Rank Adaptation (LoRA) while freezing the backbone, which significantly boosts the area under the precision‑recall curve from 0.300 to 0.679. Subsequent hard‑negative mining and supervised contrastive learning stages show negligible impact, indicating that a single LoRA adaptation is sufficient for effective underwater ATR.

By Dan Zimmerman, Frank E. Bobe III, Amelia L. McCormack, Matthew Cook, Gregory D. Vetaw
arXiv Machine Learning
Sep 25

Towards Deployable Underwater Vessel Classification

The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.

By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
arXiv Machine Learning
Jul 16

BenthiCat: An opti-acoustic dataset for advancing benthic classification and habitat mapping

arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.

By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
arXiv Machine Learning
Aug 20

Mitigating Spectral Bias in Neural Operators for Underwater Transmission Loss Prediction

The paper addresses the spectral bias of Fourier Neural Operators (FNO) in predicting underwater acoustic transmission loss. It introduces a Spectral‑Spatial Residual Learning (S2RL) framework that first uses a spectral global propagator for coarse predictions and then a spatial local refiner to recover high‑frequency details. Experiments on a South China Sea dataset show that S2RL outperforms FNO baselines while keeping inference times in the millisecond range.

By Yifan Sun, Shikai Fang, Chao Zhang, Lei Cheng, Jianlong Li, Peter Gerstoft
arXiv Computer Vision
Aug 27

RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

RSFusionDet introduces a new RGB‑Sonar multimodal object detection dataset (RSFusion) and benchmark metrics for underwater imaging. The proposed detector uses a Cross‑Attention Fusion module to align RGB and sonar features and an Object Matching Head with loss to identify identical objects across modalities. On the RSFusion dataset, RSFusionDet achieves 76.4/48.6 AP for RGB/sonar detection and 83.4 F1‑Score for cross‑modal matching, outperforming existing models and improving over the DINO baseline by 0.7/1.4 AP.

By Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang, Ye Li
arXiv AI
Jun 6

SagnacAssisted Enhanced OTDR for Distributed Acoustic Sensing: A Standardized Benchmark and Engineering Evaluation Framework

arXiv:2606. 05754v1 Announce Type: cross Abstract: Phase-sensitive optical time-domain reflectometry ($\phi$-OTDR) is widely used in large-scale distributed acoustic sensing (DAS) because it provides distributed spatiotemporal monitoring over long sensing distances.

By Weiguang Wang, Fugen Wu, Hailing Wang, Xuechen Liang, Xiaobin Li, Ru Han, Tianchang Xie
arXiv AI
Jun 12

GetNetUPAM: Ecologically Informed Nested Cross-Validation and Noise-Robust Attention for Marine Bioacoustic Monitoring

arXiv:2509. 04682v2 Announce Type: replace-cross Abstract: Deploying reliable bioacoustic monitoring systems requires models that generalize under high-noise, low-SNR conditions and evaluation protocols that expose deployment-relevant failure modes, gaps largely unaddressed in current UPAM practice.

By Nicholas R. Rasmussen, Rodrigue Rizk, Longwei Wang, KC Santosh