Hugging Face Trending Papers

Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study

Multimodal large language models (LLMs) are increasingly adopted to interpret 12-lead ECG images, though the interpretations often lack validation. However, ECG image understanding significantly differs from general images as it depends on precise waveform morphology, lead relationships and accurate interval measurements.

arXiv AI
Jul 24

Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

arXiv:2607. 20814v1 Announce Type: new Abstract: The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning models remains con- strained by limited interpretability and the hallucination risk of large language models (LLMs).

By Hai-Nam Duy Vuong, Duy-Anh Bui, Trong-Nghia Nguyen, Kim-Ngan Thi Nguyen, Trang Mai Xuan, Tien-Cuong Nguyen, Van-Dem Pham, Thien Van Luong
arXiv Computer Vision
Sep 25

A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification

The paper presents a hybrid CNN–state‑space–attention backbone designed for 12‑lead ECG classification, combining early waveform tokenization, mixed temporal dynamics modeling, and late global attention. It introduces an ECG‑oriented Joint‑Embedding Predictive Pretraining (JEPA) that samples span masks at latent resolution and predicts clean latent targets via a momentum encoder, avoiding waveform reconstruction. Experiments on CPSC2018, Chapman‑Shaoxing, and PTB‑XL, with pretraining on ~350K unlabeled CODE‑15 recordings, demonstrate strong supervised baselines and improved transfer, especially in low‑label scenarios and with LoRA adaptation.

By Yakoub Bazi, Sarah Aljuhani, Mohamad M. Al Rahhal, Mansour Zuair, Naif Alajlan
arXiv AI
Sep 21

BEAT-Net: Injecting Biomimetic Spatio-Temporal Priors for Interpretable ECG Diagnosis

BEAT-Net is a supervised biomimetic framework for ECG diagnosis that incorporates QRS-centered tokenization and a hierarchical architecture mirroring a cardiologist’s workflow. It processes heartbeat sequences through morphological, spatial, temporal, and transformer-based stages, achieving an AUC of 0.924 on large benchmarks while using only 0.7 million parameters. The model outperforms the 39.5‑million‑parameter HeartLang foundation model on morphological form classification and demonstrates superior cross‑dataset generalization with only 35% of the training data.

By Runze Ma, Haonan Lyu, Shunbo Jia, Qiang Yang, Muzi Xu, Jiaqi Zhang, Zihe Luo, Caizhi Liao
arXiv Computation and Language
Sep 1

ECGQuest: Benchmarking and Fine-Tuning Language Models for Electrocardiography

ECGQuest is a new benchmark that evaluates language models on the contextual knowledge required for electrocardiogram interpretation, featuring 10,904 True/False questions derived from 23 ECG references and 2003‑2025 Computing in Cardiology proceedings. The study tested 23 commercial and open‑source models, finding that zero‑shot accuracy ranged from 49.5% to 74.4% and that fine‑tuning with Low‑Rank Adaptation improved all open‑source models by 6.5–14.1%, with the best fine‑tuned model achieving 76.3% accuracy and a five‑model ensemble reaching 78.5%. ECGQuest demonstrates that parameter‑efficient fine‑tuning can enable smaller models to compete with larger commercial ones on ECG‑specific tasks.

By Mohammadsina Hassannia, Matthew A. Reyna, Reza Sameni
arXiv Computer Vision
Sep 24

From ECG Signals to Representative-Morphology Heatmaps for Biometric Recognition

The paper introduces representative‑morphology heatmaps, a deterministic ECG‑to‑image representation that averages the five beats closest to the block mean within each ten‑beat block, producing either a conventional trace or a dense cardiac‑time‑by‑lead heatmap. Experiments on PTB, ECG‑ID, and MIMIC‑IV‑ECG‑DEMO show that heatmaps consistently improve verification and identification performance across 15 compact models, reducing EER by an average of 9.59 percentage points and increasing Rank‑1 by 24.69 points. The study also demonstrates that ImageNet initialization benefits multilead datasets, that performance does not scale monotonically with model size, and that useful channel combinations vary by cohort and biometric task.

By Athanasios Angelakis, Marta Gomez-Barrero
arXiv AI
Aug 7

ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation

arXiv:2608. 05893v1 Announce Type: new Abstract: Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-lead ECG recordings into reliable clinical reports remains challenging.

By Akanta Das, Tasinul Islam Ahon, Ahmed Mahir Sultan Rumi, Md Mahbubur Rahman, Tausif Amim Shadly, Tanzima Hashem
arXiv Machine Learning
Aug 28

Graph-Based Pseudo-multimodal Contrastive Learning for 12-Lead ECG Representations

The paper introduces Graph-CMMC, a graph-based pseudo‑multimodal contrastive learning framework for 12‑lead ECG representations. It transforms ECG waveforms into Gramian Angular Difference Field (GADF) images to create complementary views, then aligns these views while modeling inter‑lead dependencies with a graph module. Experiments on coronary artery occlusion classification show that Graph-CMMC performs competitively with supervised methods, highlighting the value of GADF representations and explicit graph modeling for robust ECG analysis.

By Mengyu Wang, Kozo Okada, Takafumi Goto, Natsuko Jinba, Hiroki Yamaya, Kiyoshi Hibi, Tomoki Hamagami
arXiv Machine Learning
Aug 5

LAEF: A Lead-Agnostic ECG Foundation Model Towards Point-of-Care Diagnostics

arXiv:2608. 03690v1 Announce Type: new Abstract: Point-of-care cardiac devices such as smartwatches and handheld ECG recorders typically capture 1--2 leads, yet existing ECG foundation models are architecturally constrained to fixed 12-lead inputs, degrading or failing under these reduced configurations.

By Edoardo Coppola, Stefano Fiorini, Pietro Li\`o, Mattia Savardi, Alberto Signoroni