arXiv Computer Vision

A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification

The paper presents a hybrid CNN–state‑space–attention backbone designed for 12‑lead ECG classification, combining early waveform tokenization, mixed temporal dynamics modeling, and late global attention. It introduces an ECG‑oriented Joint‑Embedding Predictive Pretraining (JEPA) that samples span masks at latent resolution and predicts clean latent targets via a momentum encoder, avoiding waveform reconstruction. Experiments on CPSC2018, Chapman‑Shaoxing, and PTB‑XL, with pretraining on ~350K unlabeled CODE‑15 recordings, demonstrate strong supervised baselines and improved transfer, especially in low‑label scenarios and with LoRA adaptation.

arXiv AI
Jun 19

SL-S4Wave: Self-Supervised Learning of Physiological Waveforms with Structured State Space Models

arXiv:2606. 19888v1 Announce Type: cross Abstract: Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates, multichannel signal complexity, inherent noise, and limited labeled data.

By Feng Wu, Harsh Deep, Eric Lehman, Sanyam Kapoor, Guoshuai Zhao, Rahul Krishnan, Gari Clifford, Li-wei H Lehman
arXiv AI
Sep 21

BEAT-Net: Injecting Biomimetic Spatio-Temporal Priors for Interpretable ECG Diagnosis

BEAT-Net is a supervised biomimetic framework for ECG diagnosis that incorporates QRS-centered tokenization and a hierarchical architecture mirroring a cardiologist’s workflow. It processes heartbeat sequences through morphological, spatial, temporal, and transformer-based stages, achieving an AUC of 0.924 on large benchmarks while using only 0.7 million parameters. The model outperforms the 39.5‑million‑parameter HeartLang foundation model on morphological form classification and demonstrates superior cross‑dataset generalization with only 35% of the training data.

By Runze Ma, Haonan Lyu, Shunbo Jia, Qiang Yang, Muzi Xu, Jiaqi Zhang, Zihe Luo, Caizhi Liao
arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv Machine Learning
Jul 27

Autoregressive EHR Foundation Models with Multimodal Inputs

arXiv:2607. 22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way.

By Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
arXiv Machine Learning
Sep 10

AF-Mamba: Efficient Long-Term Signal Modeling for Early Prediction of Atrial Fibrillation Onset

AF-Mamba is a deep learning model that predicts atrial fibrillation (AF) onset one hour in advance using long‑term RR intervals. It combines temporal convolutional networks for local feature extraction with Mamba, a state‑space model for long‑range sequence modeling, achieving high sensitivity (0.889) and specificity (0.943) in subject‑wise testing. The model maintains strong performance across unseen datasets, offering a favorable trade‑off between predictive accuracy and computational efficiency for real‑time ambulatory monitoring.

By Yongbin Lee, Ki H. Chon
arXiv Machine Learning
Aug 14

CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

arXiv:2608. 12944v1 Announce Type: new Abstract: Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited.

By Hamza Shafiq, Hung Manh Pham, Bin Zhu, Pan Zhou, Jun Hu, Aaqib Saeed