arXiv:2608. 19297v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals.
By Yihan Xie, Hanwen Cui, Runze Ye, Juekai Lin, Haoyang Wang, Jinhao Mao, Bo Zhang, Wenqiao Zhang, Xiaogang Guo, Jun Xiao, Lei Zhang
arXiv:2605. 29977v2 Announce Type: replace-cross Abstract: High-fidelity ECG interpretation is increasingly reliant on massive foundation models, yet their deployment in clinical edge-care remains hindered by extreme computational demands.
By Dang Nguyen Hong, Nhi Ngoc-Yen Nguyen, Huy-Hieu Pham
ECGQuest is a new benchmark that evaluates language models on the contextual knowledge required for electrocardiogram interpretation, featuring 10,904 True/False questions derived from 23 ECG references and 2003‑2025 Computing in Cardiology proceedings. The study tested 23 commercial and open‑source models, finding that zero‑shot accuracy ranged from 49.5% to 74.4% and that fine‑tuning with Low‑Rank Adaptation improved all open‑source models by 6.5–14.1%, with the best fine‑tuned model achieving 76.3% accuracy and a five‑model ensemble reaching 78.5%. ECGQuest demonstrates that parameter‑efficient fine‑tuning can enable smaller models to compete with larger commercial ones on ECG‑specific tasks.
By Mohammadsina Hassannia, Matthew A. Reyna, Reza Sameni
The paper introduces R‑U‑Net, an ECG delineation model that combines a ResNet‑18 encoder with a U‑Net decoder. It demonstrates that this decoder design outperforms a ResNet‑18 + fully convolutional network baseline across 16 in‑domain settings and improves cross‑domain performance by 8.1 mIoU. Ablation studies reveal that the decoder contributes more to performance gains than the evaluated semi‑supervised learning methods.
By Joseph Scharpf, William Han, Chaojing Duan, Michael A. Rosenberg, Emerson Liu, Ding Zhao
arXiv:2607. 10784v1 Announce Type: cross Abstract: Deploying deep learning models for automated electrocardiogram classification on resource-constrained wearable devices remains challenging due to high computational costs.
By Yi Zhao, Jiajun Gao, Chenyang Xu, Yuxi Zhou, Hao Wang
The paper presents a hybrid CNN–state‑space–attention backbone designed for 12‑lead ECG classification, combining early waveform tokenization, mixed temporal dynamics modeling, and late global attention. It introduces an ECG‑oriented Joint‑Embedding Predictive Pretraining (JEPA) that samples span masks at latent resolution and predicts clean latent targets via a momentum encoder, avoiding waveform reconstruction. Experiments on CPSC2018, Chapman‑Shaoxing, and PTB‑XL, with pretraining on ~350K unlabeled CODE‑15 recordings, demonstrate strong supervised baselines and improved transfer, especially in low‑label scenarios and with LoRA adaptation.
By Yakoub Bazi, Sarah Aljuhani, Mohamad M. Al Rahhal, Mansour Zuair, Naif Alajlan