The paper presents a hybrid CNN–state‑space–attention backbone designed for 12‑lead ECG classification, combining early waveform tokenization, mixed temporal dynamics modeling, and late global attention. It introduces an ECG‑oriented Joint‑Embedding Predictive Pretraining (JEPA) that samples span masks at latent resolution and predicts clean latent targets via a momentum encoder, avoiding waveform reconstruction. Experiments on CPSC2018, Chapman‑Shaoxing, and PTB‑XL, with pretraining on ~350K unlabeled CODE‑15 recordings, demonstrate strong supervised baselines and improved transfer, especially in low‑label scenarios and with LoRA adaptation.
By Yakoub Bazi, Sarah Aljuhani, Mohamad M. Al Rahhal, Mansour Zuair, Naif Alajlan
arXiv:2605. 29977v2 Announce Type: replace-cross Abstract: High-fidelity ECG interpretation is increasingly reliant on massive foundation models, yet their deployment in clinical edge-care remains hindered by extreme computational demands.
By Dang Nguyen Hong, Nhi Ngoc-Yen Nguyen, Huy-Hieu Pham
arXiv:2507. 12645v1 Announce Type: cross Abstract: The increasing need for accurate and unified analysis of diverse biological signals, such as ECG and EEG, is paramount for comprehensive patient assessment, especially in synchronous monitoring.
By Mohammed Guhdar, Ramadhan J. Mstafa, Abdulhakeem O. Mohammed
arXiv:2607. 01145v2 Announce Type: replace Abstract: Data analysis in the medical domain often encounters scenarios involving a limited target dataset and a large, unannotated dataset with a general distribution.
By Siwon Kim
arXiv:2512. 13765v2 Announce Type: replace-cross Abstract: The forward problem in electrocardiology, computing body surface potentials from cardiac electrical activity, is traditionally solved using physics-based models such as the bidomain or monodomain equations.
By Shaheim Ogbomo-Harmitt, Cesare Magnetti, Chiara Spota, Jakub Grzelak, Oleg Aslanidi
arXiv:2607. 01145v1 Announce Type: new Abstract: Data analysis in the medical domain often encounters scenarios involving a limited target dataset and a large, unannotated dataset with a general distribution.
By Siwon Kim
arXiv:2606. 10802v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) typically require extensive datasets for effective training.
By Naoki Nonaka, Jun Seita
arXiv:2606. 12252v1 Announce Type: cross Abstract: Training deep neural networks for clinical time-series analysis is computationally demanding, yet many healthcare settings lack the resources required for repeated model development and deployment.
By Veerendhra Kumar Dangeti, Xiao Gu, Ying Weng, Shreyank N Gowda
BEAT-Net is a supervised biomimetic framework for ECG diagnosis that incorporates QRS-centered tokenization and a hierarchical architecture mirroring a cardiologist’s workflow. It processes heartbeat sequences through morphological, spatial, temporal, and transformer-based stages, achieving an AUC of 0.924 on large benchmarks while using only 0.7 million parameters. The model outperforms the 39.5‑million‑parameter HeartLang foundation model on morphological form classification and demonstrates superior cross‑dataset generalization with only 35% of the training data.
By Runze Ma, Haonan Lyu, Shunbo Jia, Qiang Yang, Muzi Xu, Jiaqi Zhang, Zihe Luo, Caizhi Liao
AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv:2609.08992v1 Announce Type: new
Abstract: False ventricular tachycardia (VT) alarms are a leading contributor to alarm fatigue in intensive care units. We propose a deep learning framework comb...
By Athanasios Papastathopoulos-Katsaros, Alexandra Stavrianidi, Zhandong Liu
arXiv:2608. 19297v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals.
By Yihan Xie, Hanwen Cui, Runze Ye, Juekai Lin, Haoyang Wang, Jinhao Mao, Bo Zhang, Wenqiao Zhang, Xiaogang Guo, Jun Xiao, Lei Zhang