arXiv Machine Learning

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

arXiv:2607. 27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically interpreted, or reproduced across independent analyses.

arXiv AI
Aug 5

FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection

arXiv:2608. 03597v1 Announce Type: new Abstract: Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, and mortality.

By Amirhossein Taleshinosrati, Yangyang Wang, Atitaya Phoemsuk, Vahid Abolghasemi, Naser Hossein Motlagh, Sadasivan Puthusserypady, Daniel Teichmann, Abdolrahman Peimankar
arXiv AI
Aug 7

ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation

arXiv:2608. 05893v1 Announce Type: new Abstract: Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-lead ECG recordings into reliable clinical reports remains challenging.

By Akanta Das, Tasinul Islam Ahon, Ahmed Mahir Sultan Rumi, Md Mahbubur Rahman, Tausif Amim Shadly, Tanzima Hashem
arXiv Computation and Language
Sep 1

ECGQuest: Benchmarking and Fine-Tuning Language Models for Electrocardiography

ECGQuest is a new benchmark that evaluates language models on the contextual knowledge required for electrocardiogram interpretation, featuring 10,904 True/False questions derived from 23 ECG references and 2003‑2025 Computing in Cardiology proceedings. The study tested 23 commercial and open‑source models, finding that zero‑shot accuracy ranged from 49.5% to 74.4% and that fine‑tuning with Low‑Rank Adaptation improved all open‑source models by 6.5–14.1%, with the best fine‑tuned model achieving 76.3% accuracy and a five‑model ensemble reaching 78.5%. ECGQuest demonstrates that parameter‑efficient fine‑tuning can enable smaller models to compete with larger commercial ones on ECG‑specific tasks.

By Mohammadsina Hassannia, Matthew A. Reyna, Reza Sameni
arXiv Machine Learning
Jul 21

ECG-LLM: Foundation Model for ECG-Based Cardiac Reasoning

arXiv:2607. 16323v1 Announce Type: cross Abstract: Electrocardiography (ECG) is an inexpensive, standard-of-care test for cardiac symptoms, but front-line triage often lacks immediate access to definitive imaging such as echocardiography (ECHO) or cardiac magnetic resonance (CMR).

By Alexander Selivanov, Friederike Jungmann, Jan Kehrer, Karl-Ludwig Laugwitz, Eimo Martens, Daniel Rueckert
arXiv AI
Sep 17

NeuroECG: ECGFounder-Based Deep ECG Representation for EEG-Free Neurological Prognostication After Cardiac Arrest

NeuroECG is a deep learning framework that repurposes a pretrained ECG foundation model to predict neurological outcomes after cardiac arrest without using electroencephalography (EEG). The model fine‑tunes the backbone with a gradual unfreezing strategy on single‑channel bedside ECG, aggregates multiple ECG segments via quantile pooling and PCA, and achieves an AUROC of 0.7333 using ECG alone. When combined with static clinical covariates, NeuroECG improves performance to an AUROC of 0.8077 and an AUPRC of 0.8970, demonstrating that bedside ECG can serve as a low‑cost, auxiliary prognostic tool in an EEG‑free setting.

By Jiaju Gao, Yi Zhao, Chenyang Xu, Yuxi Zhou, Hao Wang
arXiv Computer Vision
Sep 24

From ECG Signals to Representative-Morphology Heatmaps for Biometric Recognition

The paper introduces representative‑morphology heatmaps, a deterministic ECG‑to‑image representation that averages the five beats closest to the block mean within each ten‑beat block, producing either a conventional trace or a dense cardiac‑time‑by‑lead heatmap. Experiments on PTB, ECG‑ID, and MIMIC‑IV‑ECG‑DEMO show that heatmaps consistently improve verification and identification performance across 15 compact models, reducing EER by an average of 9.59 percentage points and increasing Rank‑1 by 24.69 points. The study also demonstrates that ImageNet initialization benefits multilead datasets, that performance does not scale monotonically with model size, and that useful channel combinations vary by cohort and biometric task.

By Athanasios Angelakis, Marta Gomez-Barrero
arXiv Machine Learning
Sep 7

Learning from VAE Errors to support ECG-based Differential Diagnosis of Myocardial Scar

The study investigates whether ECG representations derived from a β-variational autoencoder (VAE) can distinguish patients with myocardial scar (LGE+) from those without (LGE-) using routine ECG data. In a cohort of 300 cardiomyopathic patients, the β-VAE achieved an AUC of 0.577 (sensitivity 0.775) with Gradient Boosting, while the foundation ECGx.AI model reached an AUC of 0.686 with Random Forest. Dynamic Time Warping reconstruction errors differed significantly between classes in most leads and improved classification to an AUC of 0.643 with Logistic Regression, suggesting these errors could serve as markers of scar-related ECG changes.

By Shayan Sharifi, Riccardo Treu, Ilaria Gandin, Federico Garoia, Marco Merlo, Giulia Cisotto