AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv:2607. 16220v1 Announce Type: cross Abstract: Heart disease kills a lot of people, and one cheap way to catch it early is by listening to heart sounds with a stethoscope, or better yet, just recording them and running them through a model.
By Abhinav Pala, Dhanush Pala
arXiv:2509. 04682v2 Announce Type: replace-cross Abstract: Deploying reliable bioacoustic monitoring systems requires models that generalize under high-noise, low-SNR conditions and evaluation protocols that expose deployment-relevant failure modes, gaps largely unaddressed in current UPAM practice.
By Nicholas R. Rasmussen, Rodrigue Rizk, Longwei Wang, KC Santosh
arXiv:2606. 01483v1 Announce Type: cross Abstract: Long-form automatic speech recognition (ASR) requires both high accuracy and low latency, but existing systems force a trade-off between the two.
By Wei-Tzu Lee, Keisuke Kamahori, Baris Kasikci
arXiv:2609.22631v1 Announce Type: new
Abstract: Accurate automated interpretation of electrocardio- grams (ECGs) is essential for early detection of cardiac condi- tions such as myocardial infarction...
By Mohammad Sadman Tahsin, Haitham Y. Adarbah, Afzel Noore
BEAT-Net is a supervised biomimetic framework for ECG diagnosis that incorporates QRS-centered tokenization and a hierarchical architecture mirroring a cardiologist’s workflow. It processes heartbeat sequences through morphological, spatial, temporal, and transformer-based stages, achieving an AUC of 0.924 on large benchmarks while using only 0.7 million parameters. The model outperforms the 39.5‑million‑parameter HeartLang foundation model on morphological form classification and demonstrates superior cross‑dataset generalization with only 35% of the training data.
By Runze Ma, Haonan Lyu, Shunbo Jia, Qiang Yang, Muzi Xu, Jiaqi Zhang, Zihe Luo, Caizhi Liao
The paper presents a hybrid predictive ensemble that merges machine learning and deep neural network techniques to detect and prognosticate cardiovascular disease early. It processes real‑time physiological data from IoMT devices, applying preprocessing, feature selection, and optimized classifiers (SVM, Random Forest, XGBoost) within an ensemble architecture. The cloud‑based system achieves higher accuracy, fewer false positives, and consistent performance on real‑world datasets, supporting continuous patient monitoring and clinical decision support.
By Balaji Venkateswaran
arXiv:2606. 06718v1 Announce Type: cross Abstract: Myocardial substrate abnormalities, such as myocardial scar and myocardial infarction (MI), are associated with adverse cardiovascular outcomes.
By Canyu Lei, Fenglin Zhang, Derek Bivona, Cristiane Singulane, Jonathan Pan, Kenneth Bilchick, Amit R. Patel, Jianxin Xie
arXiv:2608.21499v1 Announce Type: cross
Abstract: Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four main auscult...
By Marcelo Nogueira, Jorge H. Oliveira, Carlos F. Ferreira, Miguel T. Coimbra, Al\'ipio M. Jorge
The paper introduces DCGCNet, a dual-codebook graph collaborative network that jointly reconstructs ECG signals and classifies atrial fibrillation. It incorporates a local‑global contrastive module for noise‑invariant feature learning and an adaptive codebook vector quantizer to prevent codebook collapse. The model achieves state‑of‑the‑art intra‑dataset performance and consistently attains AUC > 0.98 across seven cross‑dataset settings, even under realistic noisy conditions.
By Hongtao Li, Jia Wei, Guoyao Li, Yuchen Lei, Guangnian Ma, Jia Xiao, Yuanjun Lai, Shuzhen Lv, Xueqiang Ouyang
arXiv:2509. 11606v4 Announce Type: replace-cross Abstract: Cardiovascular diseases (CVDs) are the leading cause of death worldwide, accounting for approximately 17.
By Milan Marocchi, Matthew Fynn, Kayapanda Mandana, Yue Rong
The paper introduces the Speech-Unsupported Rejection Evaluation Challenge (SURE‑Challenge), a benchmark designed to test whether speech‑LLMs should accept or reject audio inputs before generating answers. Using LibriSpeech‑derived transcriptions paired with first‑word question answering, the authors evaluate various noise and silence conditions, and compare a simple energy‑plus‑Whisper‑score rule against a Qwen2‑Audio front‑end. On a 474‑row test set, the rule rejects 196 of 204 unsupported inputs while preserving accuracy on supported data, revealing a pre‑generation error mode that answer‑only scoring misses.
By Mengzhe Geng