arXiv:2608. 20315v1 Announce Type: new Abstract: Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical events.
By Jun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont, Ziv Bar-Joseph, Sven Jager, Brandon Rufino
arXiv:2606. 06990v1 Announce Type: new Abstract: The generation of high-fidelity synthetic Electronic Health Records (EHR) is crucial for advancing medical research while preserving patient privacy.
By Jalen Jiang, Chufan Gao, Ethan Rasmussen, Stephen Z. Xie, Jimeng Sun
arXiv:2511. 02340v3 Announce Type: replace Abstract: Chronic Kidney Disease (CKD) affects nearly 10\% of the global population and often progresses to end-stage renal failure.
By Yohan Lee, Dong Gyun Kang, SeHoon Park, Sa-Yoon Park, Kwangsoo Kim
MedFlow is a class‑aware multi‑scale flow matching framework designed to synthesize medical time‑series data. It uses a vector‑quantized multi‑scale tokenizer to capture both coarse and fine temporal patterns, and introduces Token Marginal Guidance to steer generation toward minority‑class characteristics. Experiments on four public datasets show MedFlow outperforms diffusion baselines, improving AUPRC by 5.8%, reducing Context‑FID by 88.6%, and achieving 3.8× higher sampling throughput.
By Yanhao Huang, Shibo Feng, Wanjin Feng, Peilin Zhao, Chunyan Miao
arXiv:2607. 06163v1 Announce Type: cross Abstract: Foundation Models for Electronic Health Records (FEMRs) are pretrained on large-scale structured patient data, enabling them to convert longitudinal patient trajectories into generalizable representations for diverse clinical prediction tasks.
By Jie Huang, Pengfei Yin, Zihan Xu, Daniel Capurro, Mike Conway, Ting Dang
arXiv:2608. 06265v1 Announce Type: new Abstract: Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access.
By Omid Bazgir, Md Nasir, Jacob Hoffman, Yang Yang, Manu Agrawal, Anusua Trivedi, Vinay Rao Dandin, Chris Gibbons, Christine Swisher
The study benchmarks tokenization choices for generative medical event models, evaluating quantization granularity, reference-range anchoring, code–value fusion, numeric and temporal encodings, and native versus harmonized event representations. Using Llama and Qwen architectures, 156 models were trained and assessed on early hospitalization data, showing that fusing codes with value deciles and using event-order or admission-relative RoPE embeddings improved predictive performance. The Common Longitudinal Intensive Care Unit Data Format (CLIF) reduced token count by 30.8% while enhancing outcomes in most families.
By Inhyeok Lee, Luke Solo, Michael C. Burkhart, Bashar Ramadan, Sahil Sethi, Sarah Jabbour, William F. Parker, Brett K. Beaulieu-Jones
The paper introduces Scaling Electronic Health Record Foundation Models for Population Health Management, a large‑scale model trained on billions of medical events from over 5 million patients in Taiwan and the United States. By aligning ICD codes across different health systems, the model achieves strong scaling and generalization across 11 chronic disease prediction tasks, outperforming tree‑based, general, and biomedical language models with high sensitivity at 99% specificity. It also demonstrates superior few‑shot performance on the EHRShot benchmark and shows that cross‑system alignment provides a stronger pretraining signal than single‑site duplication in data‑limited scenarios.
By Liwen Sun, Hao-Ren Yao, Ophir Frieder, Xiang Qian, Chenyan Xiong
arXiv:2605. 30295v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in realistic, electronic health record-congruent settings remains limited.
By Valentina Bui Muti, Eug\'enie Dulout, Ziquan Fu
arXiv:2608. 12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identification.
By Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto, David Rehkopf, Ayin Vala, Tanmoy Sarkar Pias
arXiv:2509. 24118v2 Announce Type: replace Abstract: Electronic health Records (EHRs) have become a cornerstone in modern-day healthcare.
By Md Mozaharul Mottalib, Thao-Ly T. Phan, Rahmatollah Beheshti
The study re‑implements 12 AI algorithms for electronic health records within a unified framework and evaluates them on MIMIC‑IV and NWICU datasets. It compares expert‑authored clinically meaningful tasks with randomly generated tasks, finding that pairwise algorithm comparisons transfer well across task families and datasets, yet clinically meaningful tasks show stronger task‑method interactions. The results also reveal that newer algorithms do not consistently outperform older ones, with gradient‑boosted trees remaining highly competitive when combined with modern EHR representations.
By Florent Pollet, Matthew McDermott