arXiv AI

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

arXiv AI
Jun 19

SL-S4Wave: Self-Supervised Learning of Physiological Waveforms with Structured State Space Models

arXiv:2606. 19888v1 Announce Type: cross Abstract: Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates, multichannel signal complexity, inherent noise, and limited labeled data.

By Feng Wu, Harsh Deep, Eric Lehman, Sanyam Kapoor, Guoshuai Zhao, Rahul Krishnan, Gari Clifford, Li-wei H Lehman
arXiv AI
Jun 2

DAStatFormer: A Hybrid Multibranch Transformer with Statistical Feature Integration for DAS-Based Pattern Recognitions

arXiv:2606. 00081v1 Announce Type: cross Abstract: Distributed Acoustic Sensing (DAS) enables large-scale monitoring through optical fibers, but its high dimensionality and complex spatio-temporal patterns make event classification demanding.

By Michel Dione (CERI SN - IMT Nord Europe), Jerry Lonlac (CERI SN - IMT Nord Europe), H\'el\`ene Louis (CERI SN - IMT Nord Europe), Anthony Fleury (CERI SN - IMT Nord Europe), Stephane Lecoeuche
arXiv AI
Sep 1

Frequency Selective Neural Networks as a Foundation Architecture for Time Series Learning

The paper introduces the Frequency Selective Neural Network (FSNN), a new foundation architecture for time‑series learning that embeds advanced signal‑processing mathematics into its neural topology. By using a fully differentiable Wiener‑like filter bank optimized with complex‑domain backpropagation, FSNN autonomously discovers and isolates the precise physical modes of a given task, thereby avoiding the spectral entanglement that plagues CNNs, RNNs, and Transformers. Extensive evaluations show that FSNN achieves state‑of‑the‑art predictive performance, attaining 77.0 % average accuracy on the 10 multivariate UEA datasets and leading all major metrics on the imbalanced PTB‑XL ECG benchmark, while converging directly on physically meaningful frequency bands such as the cardiac QRS complex.

By Hui Huang, Ye Sun, Shiyan Hu
arXiv Machine Learning
Sep 10

AF-Mamba: Efficient Long-Term Signal Modeling for Early Prediction of Atrial Fibrillation Onset

AF-Mamba is a deep learning model that predicts atrial fibrillation (AF) onset one hour in advance using long‑term RR intervals. It combines temporal convolutional networks for local feature extraction with Mamba, a state‑space model for long‑range sequence modeling, achieving high sensitivity (0.889) and specificity (0.943) in subject‑wise testing. The model maintains strong performance across unseen datasets, offering a favorable trade‑off between predictive accuracy and computational efficiency for real‑time ambulatory monitoring.

By Yongbin Lee, Ki H. Chon
arXiv AI
Sep 10

S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network

The paper introduces S$^3$F-Net, a dual‑branch network that fuses spatial and spectral representations for medical image classification. It combines a deep spatial CNN with a shallow spectral encoder, SpectraNet, which uses a learnable SpectralFilter layer to process the full Fourier spectrum efficiently. Evaluated on four medical imaging datasets, S$^3$F-Net consistently outperforms spatial‑only baselines, achieving state‑of‑the‑art accuracy on BRISC2025 and surpassing deeper models on the Chest X‑Ray Pneumonia dataset.

By Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan
arXiv AI
Sep 2

MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries

MADS (Multi-view Acoustic Descriptor Set) is a compact 19‑dimensional, physics‑informed descriptor set designed to capture spectral, temporal, mechanical, and stochastic aspects of audio signals. Unlike traditional log‑mel or MFCC representations, MADS encodes excitation, damping, periodicity, impulsiveness, and structural consistency in a unified multi‑view format. Evaluated on ESC‑10, ESC‑50, and MSoS datasets with classical machine learning models, MADS outperforms conventional 26‑D MFCC and 38‑D spectral‑summary baselines, achieving 81.00% on ESC‑10, 52.78% on ESC‑50, and 67.48% on MSoS while using roughly half the dimensionality of the 38‑D baseline.

By Utsab Ghosh, Roshni Chakraborty