arXiv Machine Learning

Linear probing enables Ship-Radiated Noise recognition with pretrained audio embeddings

arXiv Machine Learning
Sep 25

Towards Deployable Underwater Vessel Classification

The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.

By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
arXiv Machine Learning
Sep 15

Large-scale bioacoustic detection using semantic segmentation: a deep learning framework applied to fin whale calls in ocean-bottom seismometer recordings

arXiv:2609.13281v1 Announce Type: cross Abstract: Ocean-bottom seismometers (OBS), originally deployed for geophysical research, continuously record low-frequency sound for months to years across bro...

By Jocelyn Japnanto, Alex A. Saoulis, Miriam Romagosa, Rita Leit\~ao, Gabrielle Arrieta, M\'onica A. Silva, Matthew Graham, Ana M. G. Ferreira
arXiv Computer Vision
Sep 3

Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade

The paper presents a lightweight ResNet-based two-stage cascade for passive acoustic monitoring of killer whales. First, it detects vocalizations, then it classifies confident detections into five eastern North Pacific ecotypes, abstaining on ambiguous calls. The pipeline achieves high macro‑F1 scores on the DCLDE 2027 dataset and improves real‑time inference speed, while active learning adapts the detector to new acoustic environments.

By Daniela Ruiz, Manuel Castellote, Zhongqi Miao, Carl Chalmers, Bruno Demuro, Rahul Dodhia, Pablo Arbelaez, Juan M. Lavista
arXiv AI
Sep 21

LoRA Enhanced Contrastive Learning with SAS Vision Transformers

The paper presents a three‑stage, parameter‑efficient approach to improve automatic target recognition (ATR) with synthetic aperture sonar (SAS) data by adapting DINOv3 Vision Transformers. Stage 1 applies Low‑Rank Adaptation (LoRA) while freezing the backbone, which significantly boosts the area under the precision‑recall curve from 0.300 to 0.679. Subsequent hard‑negative mining and supervised contrastive learning stages show negligible impact, indicating that a single LoRA adaptation is sufficient for effective underwater ATR.

By Dan Zimmerman, Frank E. Bobe III, Amelia L. McCormack, Matthew Cook, Gregory D. Vetaw