arXiv Computer Vision

ECG-Mamba-V2: Architectural Refinements to a Bidirectional State Space Model for Multi-Label 12-Lead ECG Classification

ECG-Mamba-V2 refines a bidirectional Vision Mamba encoder for multi‑label 12‑lead ECG classification by appending the class token at the sequence end, summing forward and backward outputs without ½ scaling, and applying uniform dropout across blocks. These architectural tweaks yield a 0.6494 macro AUPRC and 0.9716 macro AUROC on the PhysioNet/CinC Challenge 2021, outperforming its predecessor while using 34% fewer parameters and achieving 38% higher throughput.

arXiv AI
Sep 15

On the role of the tokenizer in ECG transformer models

The study investigates how different tokenization methods affect ECG Transformer models by comparing eight strategies across four backbone architectures on the CPSC2018 classification task. Physiology-aware tokenizations such as Median-beat and HeartLang achieve higher mean macro-AUCs (0.893 and 0.889) than point-wise and patch-wise approaches (0.822 and 0.824), while also reducing sequence length and training memory usage. Combining the two physiology-aware representations further improves macro-AUC by 8.2%.

By Jiawei Li, Fabio Bonassi, Johan Sundstr\"om, Thomas B. Sch\"on, Ant\^onio H. Ribeiro
arXiv Computer Vision
Sep 25

A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification

The paper presents a hybrid CNN–state‑space–attention backbone designed for 12‑lead ECG classification, combining early waveform tokenization, mixed temporal dynamics modeling, and late global attention. It introduces an ECG‑oriented Joint‑Embedding Predictive Pretraining (JEPA) that samples span masks at latent resolution and predicts clean latent targets via a momentum encoder, avoiding waveform reconstruction. Experiments on CPSC2018, Chapman‑Shaoxing, and PTB‑XL, with pretraining on ~350K unlabeled CODE‑15 recordings, demonstrate strong supervised baselines and improved transfer, especially in low‑label scenarios and with LoRA adaptation.

By Yakoub Bazi, Sarah Aljuhani, Mohamad M. Al Rahhal, Mansour Zuair, Naif Alajlan
arXiv AI
Jun 8

LuMamba: Latent Unified Mamba for Electrode Topology-Invariant and Efficient EEG Modeling

arXiv:2603. 19100v2 Announce Type: replace Abstract: Electroencephalography (EEG) enables non-invasive monitoring of brain activity across clinical and neurotechnology applications, yet building foundation models for EEG remains challenging due to differing electrode topologies and computational scalability, as Transformer architectures incur quadratic sequence complexity.

By Dana\'e Broustail, Anna Tegon, Thorir Mar Ingolfsson, Yawei Li, Luca Benini
arXiv AI
Oct 1

Robust Transfer Learning for Paper ECG Recognition

RobECG-CL is a rank‑aware contrastive learning framework designed to learn robust representations for paper ECG images, which often suffer from varied layouts, physical artifacts, and limited labels. The method constructs progressively degraded views from standard 12‑lead ECG recordings and trains the model to maintain invariance to the same recording while respecting degradation ordering. In synthetic stress tests on CODE‑II and EchoNext, RobECG‑CL shows improved robustness under severe degradation and few‑shot transfer, outperforming other contrastive baselines and surpassing the waveform‑based foundation model ECG‑FM in a 1% labeled setting; it also achieves the best macro AUROC on 312 hospital samples with 37 labels.

By Yinghao Xie, Zhenbang Dai, Haojun Wang, Jinyu Cai, Fabio Bonassi, Hongwu Chen, Johan Sundstr\"om, Jiawei Li, Ant\^onio H. Ribeiro
arXiv Machine Learning
Sep 1

Beat-Synchronous Tokenization for ECG Transformers

arXiv:2608.30367v1 Announce Type: new Abstract: Transformer-based electrocardiogram (ECG) models commonly tokenize waveforms into fixed temporal patches. Though convenient, fixed patching can split h...

By Ahmed Sameh, Nolan Wilson, Max Enderlein, Yogatheesan Varatharajah
arXiv AI
Aug 5

FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection

arXiv:2608. 03597v1 Announce Type: new Abstract: Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, and mortality.

By Amirhossein Taleshinosrati, Yangyang Wang, Atitaya Phoemsuk, Vahid Abolghasemi, Naser Hossein Motlagh, Sadasivan Puthusserypady, Daniel Teichmann, Abdolrahman Peimankar
arXiv AI
Sep 16

Decoder Design Matters for ECG Delineation

The paper introduces R‑U‑Net, an ECG delineation model that combines a ResNet‑18 encoder with a U‑Net decoder. It demonstrates that this decoder design outperforms a ResNet‑18 + fully convolutional network baseline across 16 in‑domain settings and improves cross‑domain performance by 8.1 mIoU. Ablation studies reveal that the decoder contributes more to performance gains than the evaluated semi‑supervised learning methods.

By Joseph Scharpf, William Han, Chaojing Duan, Michael A. Rosenberg, Emerson Liu, Ding Zhao
arXiv AI
Jul 7

ImputeECG: Deep Learning Reconstruction of Complete 12-Lead Electrocardiograms from Incomplete Recordings for Cardiac Assessment

arXiv:2607. 05009v1 Announce Type: cross Abstract: Complete digital 12-lead electrocardiograms (ECGs) are essential for AI-enabled cardiovascular assessment, yet many clinical ECG records, particularly those digitized from ECG images, remain incomplete because of short display formats, incomplete waveform digitization, lead loss, or signal corruption.

By Xiaocheng Fang, Haoyu Wang, Jieyi Cai, Qinghao Zhao, Jun Li, Shanwei Zhang, Guangkun Nie, Yujie Xiao, Shun Huang, Jiarui Jin, Hongmin Liu, Guodong Wang, Shuohua Chen, Liming Lin, Shouling Wu, Hongyan Li, Shenda Hong
arXiv Machine Learning
Jul 31

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

arXiv:2607. 27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically interpreted, or reproduced across independent analyses.

By Yixuan Duan, Wei Qiu