arXiv Computation and Language By Mohammadsina Hassannia, Matthew A. Reyna, Reza Sameni

ECGQuest: Benchmarking and Fine-Tuning Language Models for Electrocardiography

Read the original on arXiv Computation and Language →

ECGQuest is a new benchmark that evaluates language models on the contextual knowledge required for electrocardiogram interpretation, featuring 10,904 True/False questions derived from 23 ECG references and 2003‑2025 Computing in Cardiology proceedings. The study tested 23 commercial and open‑source models, finding that zero‑shot accuracy ranged from 49.5% to 74.4% and that fine‑tuning with Low‑Rank Adaptation improved all open‑source models by 6.5–14.1%, with the best fine‑tuned model achieving 76.3% accuracy and a five‑model ensemble reaching 78.5%. ECGQuest demonstrates that parameter‑efficient fine‑tuning can enable smaller models to compete with larger commercial ones on ECG‑specific tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jul 21

ECG-LLM: Foundation Model for ECG-Based Cardiac Reasoning

arXiv:2607. 16323v1 Announce Type: cross Abstract: Electrocardiography (ECG) is an inexpensive, standard-of-care test for cardiac symptoms, but front-line triage often lacks immediate access to definitive imaging such as echocardiography (ECHO) or cardiac magnetic resonance (CMR).

By Alexander Selivanov, Friederike Jungmann, Jan Kehrer, Karl-Ludwig Laugwitz, Eimo Martens, Daniel Rueckert
arXiv AI
Jul 24

Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

arXiv:2607. 20814v1 Announce Type: new Abstract: The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning models remains con- strained by limited interpretability and the hallucination risk of large language models (LLMs).

By Hai-Nam Duy Vuong, Duy-Anh Bui, Trong-Nghia Nguyen, Kim-Ngan Thi Nguyen, Trang Mai Xuan, Tien-Cuong Nguyen, Van-Dem Pham, Thien Van Luong
arXiv AI
Aug 7

ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation

arXiv:2608. 05893v1 Announce Type: new Abstract: Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-lead ECG recordings into reliable clinical reports remains challenging.

By Akanta Das, Tasinul Islam Ahon, Ahmed Mahir Sultan Rumi, Md Mahbubur Rahman, Tausif Amim Shadly, Tanzima Hashem
arXiv Machine Learning
Aug 11

Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reasoning Across Diseases and Populations

arXiv:2608. 09053v1 Announce Type: cross Abstract: Cardiologists interpret electrocardiograms by localizing waveform components, measuring rhythm and interval patterns, and translating these structured observations into diagnostic evidence.

By Hongxiang Gao, He-yang Xu, Yuwen Li, Minghui Zhao, Zhipeng Cai, Xingyao Wang, Chenxi Yang, Jianqing Li, Chengyu Liu
arXiv Machine Learning
Jul 31

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

arXiv:2607. 27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically interpreted, or reproduced across independent analyses.

By Yixuan Duan, Wei Qiu