arXiv Machine Learning

The Vote Hides the Failure: Aggregation Choice and Noise Robustness in Heart Murmur Detection

arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv AI
Jun 12

GetNetUPAM: Ecologically Informed Nested Cross-Validation and Noise-Robust Attention for Marine Bioacoustic Monitoring

arXiv:2509. 04682v2 Announce Type: replace-cross Abstract: Deploying reliable bioacoustic monitoring systems requires models that generalize under high-noise, low-SNR conditions and evaluation protocols that expose deployment-relevant failure modes, gaps largely unaddressed in current UPAM practice.

By Nicholas R. Rasmussen, Rodrigue Rizk, Longwei Wang, KC Santosh
arXiv AI
Sep 21

BEAT-Net: Injecting Biomimetic Spatio-Temporal Priors for Interpretable ECG Diagnosis

BEAT-Net is a supervised biomimetic framework for ECG diagnosis that incorporates QRS-centered tokenization and a hierarchical architecture mirroring a cardiologist’s workflow. It processes heartbeat sequences through morphological, spatial, temporal, and transformer-based stages, achieving an AUC of 0.924 on large benchmarks while using only 0.7 million parameters. The model outperforms the 39.5‑million‑parameter HeartLang foundation model on morphological form classification and demonstrates superior cross‑dataset generalization with only 35% of the training data.

By Runze Ma, Haonan Lyu, Shunbo Jia, Qiang Yang, Muzi Xu, Jiaqi Zhang, Zihe Luo, Caizhi Liao
arXiv AI
Sep 7

A Hybrid Predictive Ensemble of Machine Learning and Deep Neural Networks for Early Cardiovascular Disease Risk Assessment

The paper presents a hybrid predictive ensemble that merges machine learning and deep neural network techniques to detect and prognosticate cardiovascular disease early. It processes real‑time physiological data from IoMT devices, applying preprocessing, feature selection, and optimized classifiers (SVM, Random Forest, XGBoost) within an ensemble architecture. The cloud‑based system achieves higher accuracy, fewer false positives, and consistent performance on real‑world datasets, supporting continuous patient monitoring and clinical decision support.

By Balaji Venkateswaran
arXiv Machine Learning
Aug 20

Atrial Fibrillation Detection with Arbitrary Leads via a Codebook-Based Reconstruction-Classification Framework

The paper introduces DCGCNet, a dual-codebook graph collaborative network that jointly reconstructs ECG signals and classifies atrial fibrillation. It incorporates a local‑global contrastive module for noise‑invariant feature learning and an adaptive codebook vector quantizer to prevent codebook collapse. The model achieves state‑of‑the‑art intra‑dataset performance and consistently attains AUC > 0.98 across seven cross‑dataset settings, even under realistic noisy conditions.

By Hongtao Li, Jia Wei, Guoyao Li, Yuchen Lei, Guangnian Ma, Jia Xiao, Yuanjun Lai, Shuzhen Lv, Xueqiang Ouyang
arXiv Computation and Language
Aug 31

SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation

The paper introduces the Speech-Unsupported Rejection Evaluation Challenge (SURE‑Challenge), a benchmark designed to test whether speech‑LLMs should accept or reject audio inputs before generating answers. Using LibriSpeech‑derived transcriptions paired with first‑word question answering, the authors evaluate various noise and silence conditions, and compare a simple energy‑plus‑Whisper‑score rule against a Qwen2‑Audio front‑end. On a 474‑row test set, the rule rejects 196 of 204 unsupported inputs while preserving accuracy on supported data, revealing a pre‑generation error mode that answer‑only scoring misses.

By Mengzhe Geng