arXiv Computation and Language By Liwen Sun, Hao-Ren Yao, Ophir Frieder, Xiang Qian, Chenyan Xiong

Scaling Electronic Health Record Foundation Models for Population Health Management

Read the original on arXiv Computation and Language →

The paper introduces Scaling Electronic Health Record Foundation Models for Population Health Management, a large‑scale model trained on billions of medical events from over 5 million patients in Taiwan and the United States. By aligning ICD codes across different health systems, the model achieves strong scaling and generalization across 11 chronic disease prediction tasks, outperforming tree‑based, general, and biomedical language models with high sensitivity at 99% specificity. It also demonstrates superior few‑shot performance on the EHRShot benchmark and shows that cross‑system alignment provides a stronger pretraining signal than single‑site duplication in data‑limited scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 18

Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic

arXiv:2608. 16273v1 Announce Type: cross Abstract: Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research.

By Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson
arXiv AI
Sep 3

General Demographic Pre-trained Models for Enhancing Predictive Performance Across Diseases and Population

The paper introduces the General Demographic Pre-trained (GDP) model, a lightweight foundation model that learns representations from the two most common clinical attributes—age and sex. By optimizing encoding and visit‑reordering strategies, GDP embeddings are shown to improve predictive performance when concatenated with raw features across various disease and geographic cohorts. The model outperforms several state‑of‑the‑art tabular foundation models and tree‑based algorithms, demonstrating that enriched demographic embeddings can enhance classification tasks while remaining fully compatible with standard classifiers.

By Li-Chin Chen, Ji-Tian Sheu, Yuh-Jue Chuang
arXiv AI
Jun 16

Hierarchical Modeling of ICD Codes in EHR Foundation Models

arXiv:2606. 15447v1 Announce Type: new Abstract: Electronic health record foundation models typically treat ICD diagnosis codes as flat tokens, overlooking the clinically meaningful hierarchical structure that captures disease families, subcategories, and fine-grained diagnostic detail.

By Megha Thukral, Dong Gyun Kang, Rudra Pratap Singh, Shruthi Kashinath Hiremath, Katrin H\"ansel, Thomas Pl\"otz
arXiv AI
Aug 11

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

arXiv:2505. 16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability.

By Vincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang, Florent Pollet, Young Sang Choi, Xinzhuo Jiang, Yuta Kobayashi, Yanwei Li, Sara Matijevic, Karthik Natarajan, Shalmali Joshi
arXiv Machine Learning
Aug 21

Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records

arXiv:2608. 20315v1 Announce Type: new Abstract: Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical events.

By Jun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont, Ziv Bar-Joseph, Sven Jager, Brandon Rufino