arXiv AI

Analysis of the Shortcut Learning and Clever Hans Effect in CNN based ECG Image Classification

arXiv:2607. 25117v1 Announce Type: cross Abstract: Deep learning models for ECG image classification may achieve high accuracy by exploiting non-physiological visual cues instead of ECG waveform morphology.

arXiv AI
Jul 24

Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

arXiv:2607. 20814v1 Announce Type: new Abstract: The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning models remains con- strained by limited interpretability and the hallucination risk of large language models (LLMs).

By Hai-Nam Duy Vuong, Duy-Anh Bui, Trong-Nghia Nguyen, Kim-Ngan Thi Nguyen, Trang Mai Xuan, Tien-Cuong Nguyen, Van-Dem Pham, Thien Van Luong
arXiv Machine Learning
Sep 15

Bridging the Gap in ECG-Based Emotion Recognition: A Unified Evaluation of Deep Learning Models

The paper evaluates deep learning models for electrocardiogram‑based emotion recognition, focusing on generalization across datasets rather than dataset‑specific performance. It introduces two open‑source tools—ARRC for standardized benchmarking and ARDT for inter‑dataset training—to merge three public AER datasets (CUADS, ASCERTAIN, DREAMER) into a more variable benchmark. Using these tools, the authors compare three prominent deep learning architectures and two CNN baselines with hyperparameter tuning and 10‑fold cross‑validation, revealing trade‑offs between accuracy and model complexity and providing a reproducible benchmark for future research.

By Timothy C Sweeney-Fanelli, Ajan Ahmed, Masudul Imtiaz
Hugging Face Trending Papers
Jun 22

Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study

Multimodal large language models (LLMs) are increasingly adopted to interpret 12-lead ECG images, though the interpretations often lack validation. However, ECG image understanding significantly differs from general images as it depends on precise waveform morphology, lead relationships and accurate interval measurements.

arXiv Machine Learning
Sep 14

A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients

The paper presents a new labelled ICU dataset and benchmarks for detecting atrial fibrillation (AF) from electrocardiograms (ECGs). It compares three AI approaches—feature‑based classifiers, deep learning, and ECG foundation models—across Canadian ICU data and the 2021 PhysioNet challenge, finding that ECG foundation models with transfer learning achieve the highest F1 score (0.89). The study demonstrates the feasibility of automated AF monitoring in ICU settings and provides resources for further research.

By Sarah Nassar, Nooshin Maghsoodi, Sophia Mannina, Shamel Addas, Stephanie Sibley, Gabor Fichtinger, David Pichora, David Maslove, Purang Abolmaesumi, Parvin Mousavi
arXiv Computer Vision
Sep 24

From ECG Signals to Representative-Morphology Heatmaps for Biometric Recognition

The paper introduces representative‑morphology heatmaps, a deterministic ECG‑to‑image representation that averages the five beats closest to the block mean within each ten‑beat block, producing either a conventional trace or a dense cardiac‑time‑by‑lead heatmap. Experiments on PTB, ECG‑ID, and MIMIC‑IV‑ECG‑DEMO show that heatmaps consistently improve verification and identification performance across 15 compact models, reducing EER by an average of 9.59 percentage points and increasing Rank‑1 by 24.69 points. The study also demonstrates that ImageNet initialization benefits multilead datasets, that performance does not scale monotonically with model size, and that useful channel combinations vary by cohort and biometric task.

By Athanasios Angelakis, Marta Gomez-Barrero
arXiv Computer Vision
Sep 25

A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification

The paper presents a hybrid CNN–state‑space–attention backbone designed for 12‑lead ECG classification, combining early waveform tokenization, mixed temporal dynamics modeling, and late global attention. It introduces an ECG‑oriented Joint‑Embedding Predictive Pretraining (JEPA) that samples span masks at latent resolution and predicts clean latent targets via a momentum encoder, avoiding waveform reconstruction. Experiments on CPSC2018, Chapman‑Shaoxing, and PTB‑XL, with pretraining on ~350K unlabeled CODE‑15 recordings, demonstrate strong supervised baselines and improved transfer, especially in low‑label scenarios and with LoRA adaptation.

By Yakoub Bazi, Sarah Aljuhani, Mohamad M. Al Rahhal, Mansour Zuair, Naif Alajlan