arXiv AI By Jiawei Li, Fabio Bonassi, Johan Sundstr\"om, Thomas B. Sch\"on, Ant\^onio H. Ribeiro

On the role of the tokenizer in ECG transformer models

Read the original on arXiv AI →

The study investigates how different tokenization methods affect ECG Transformer models by comparing eight strategies across four backbone architectures on the CPSC2018 classification task. Physiology-aware tokenizations such as Median-beat and HeartLang achieve higher mean macro-AUCs (0.893 and 0.889) than point-wise and patch-wise approaches (0.822 and 0.824), while also reducing sequence length and training memory usage. Combining the two physiology-aware representations further improves macro-AUC by 8.2%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 1

Beat-Synchronous Tokenization for ECG Transformers

arXiv:2608.30367v1 Announce Type: new Abstract: Transformer-based electrocardiogram (ECG) models commonly tokenize waveforms into fixed temporal patches. Though convenient, fixed patching can split h...

By Ahmed Sameh, Nolan Wilson, Max Enderlein, Yogatheesan Varatharajah
arXiv Machine Learning
Jul 27

Autoregressive EHR Foundation Models with Multimodal Inputs

arXiv:2607. 22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way.

By Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
arXiv Machine Learning
Jul 31

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

arXiv:2607. 27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically interpreted, or reproduced across independent analyses.

By Yixuan Duan, Wei Qiu