arXiv Machine Learning

Towards Deployable Underwater Vessel Classification

The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.

arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv Computer Vision
Sep 3

Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks

The paper investigates Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) imagery, comparing modern convolutional neural networks (CNNs) and transformer-based deep neural networks (DNNs). It examines how factors such as network size, architecture, pretraining methods, data augmentation, and regularization influence performance, aiming to identify the highest-performing model and provide a training roadmap for state‑of‑the‑art SAS‑ATR systems.

By C. J. Moore, Alex Hurt, Jordan Malof
arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv Machine Learning
Sep 15

Large-scale bioacoustic detection using semantic segmentation: a deep learning framework applied to fin whale calls in ocean-bottom seismometer recordings

arXiv:2609.13281v1 Announce Type: cross Abstract: Ocean-bottom seismometers (OBS), originally deployed for geophysical research, continuously record low-frequency sound for months to years across bro...

By Jocelyn Japnanto, Alex A. Saoulis, Miriam Romagosa, Rita Leit\~ao, Gabrielle Arrieta, M\'onica A. Silva, Matthew Graham, Ana M. G. Ferreira
arXiv Computer Vision
Sep 3

Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade

The paper presents a lightweight ResNet-based two-stage cascade for passive acoustic monitoring of killer whales. First, it detects vocalizations, then it classifies confident detections into five eastern North Pacific ecotypes, abstaining on ambiguous calls. The pipeline achieves high macro‑F1 scores on the DCLDE 2027 dataset and improves real‑time inference speed, while active learning adapts the detector to new acoustic environments.

By Daniela Ruiz, Manuel Castellote, Zhongqi Miao, Carl Chalmers, Bruno Demuro, Rahul Dodhia, Pablo Arbelaez, Juan M. Lavista
arXiv AI
Jun 2

DAStatFormer: A Hybrid Multibranch Transformer with Statistical Feature Integration for DAS-Based Pattern Recognitions

arXiv:2606. 00081v1 Announce Type: cross Abstract: Distributed Acoustic Sensing (DAS) enables large-scale monitoring through optical fibers, but its high dimensionality and complex spatio-temporal patterns make event classification demanding.

By Michel Dione (CERI SN - IMT Nord Europe), Jerry Lonlac (CERI SN - IMT Nord Europe), H\'el\`ene Louis (CERI SN - IMT Nord Europe), Anthony Fleury (CERI SN - IMT Nord Europe), Stephane Lecoeuche