SeisBench DAS is an extension of the SeisBench library that standardizes distributed acoustic sensing (DAS) data, metadata, labels, and models for machine learning. It leverages the xdas framework for data ingestion and virtual array handling, and PyTorch for model application, providing an efficient engine to apply deep learning models across diverse DAS formats. The framework aims to bridge the gap between model developers and practitioners, facilitating the adoption of deep learning in DAS research and allowing easy integration of future developments.
By Jannes M\"unchmeyer, Han Xiao, Frederik Tilmann
AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv:2606. 05754v1 Announce Type: cross Abstract: Phase-sensitive optical time-domain reflectometry ($\phi$-OTDR) is widely used in large-scale distributed acoustic sensing (DAS) because it provides distributed spatiotemporal monitoring over long sensing distances.
By Weiguang Wang, Fugen Wu, Hailing Wang, Xuechen Liang, Xiaobin Li, Ru Han, Tianchang Xie
arXiv:2609.22139v1 Announce Type: cross
Abstract: Automatic modulation classification (AMC) of received radio signals is prudent for further signal processing tasks such as communication monitoring,...
By Qamar Ijaz, Nayyer Aafaq
arXiv:2606. 11922v1 Announce Type: cross Abstract: Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Transformer (AST).
By Hemansh Shridhar, Miika Toikkanen, June-Woo Kim
arXiv:2606. 19888v1 Announce Type: cross Abstract: Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates, multichannel signal complexity, inherent noise, and limited labeled data.
By Feng Wu, Harsh Deep, Eric Lehman, Sanyam Kapoor, Guoshuai Zhao, Rahul Krishnan, Gari Clifford, Li-wei H Lehman
The paper introduces m-WCN, an end‑to‑end deep learning framework that neuralizes multi‑wavelet decomposition to jointly extract temporal patterns and frequency components from time series. Two task‑specific architectures built on m‑WCN—TFBC for classification and FTB for forecasting—are shown to outperform baseline models on 64 UCR datasets and seven forecasting benchmarks, achieving average improvements of nearly 20% in both tasks. The approach leverages trainable convolutional operators and orthogonality constraints to produce interpretable multi‑resolution representations.
By Xiaohan Jiang, Jingyuan Wang, Jiahao Ji, Yongyao Wang, Chen Yang, Junjie Wu
arXiv:2606. 02341v1 Announce Type: cross Abstract: Underwater acoustic classification has a wide array of oceanic applications, but faces challenges due to an increasingly complex acoustic environment.
By Amirmohammad Mohammadi, Joshua Peeples, Alexandra Van Dine
The paper presents a method that uses distributed acoustic sensing (DAS) and deep learning to monitor urban traffic at fine spatiotemporal resolution. By repurposing underground fiber‑optic cables as dense sensor arrays, the approach captures roadway activity at meter‑level spatial and second‑level temporal scales. A deep learning framework processes raw vibration waveforms to detect vehicle trajectories and infer traffic volume and speed, with a hybrid training strategy that combines synthetic and manually annotated data to improve detection in noisy, congested conditions.
By Hao Tian, Heng Cai, Xiaowei Chen, Yifan Yang
arXiv:2606. 14120v1 Announce Type: cross Abstract: Auditory attention decoding (AAD) aims to infer the attended speaker from neural responses in multi-speaker acoustic environments and is a key problem for neuro-steered hearing systems.
By Ziwei Wang, Xingyi He, Tianwang Jia, Hongbin Wang, Dongrui Wu
The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.
By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
MADS (Multi-view Acoustic Descriptor Set) is a compact 19‑dimensional, physics‑informed descriptor set designed to capture spectral, temporal, mechanical, and stochastic aspects of audio signals. Unlike traditional log‑mel or MFCC representations, MADS encodes excitation, damping, periodicity, impulsiveness, and structural consistency in a unified multi‑view format. Evaluated on ESC‑10, ESC‑50, and MSoS datasets with classical machine learning models, MADS outperforms conventional 26‑D MFCC and 38‑D spectral‑summary baselines, achieving 81.00% on ESC‑10, 52.78% on ESC‑50, and 67.48% on MSoS while using roughly half the dimensionality of the 38‑D baseline.
By Utsab Ghosh, Roshni Chakraborty