Hugging Face Trending Papers

AdaPCLA: Adaptive Prior-Calibrated Logit Adjustment for Long-Tailed Longitudinal EHR Generation

Generative modeling of longitudinal Electronic Health Records is increasingly important for privacy-preserving research, yet standard autoregressive models tend to underrepresent the co-occurrence structure of tail events (i. e.

arXiv AI
Jul 23

SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework

arXiv:2607. 19524v1 Announce Type: cross Abstract: Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks.

By Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown
arXiv Machine Learning
Jul 27

Pretraining EHR Foundation Models with Patient-Aware Sampling

arXiv:2607. 22114v1 Announce Type: new Abstract: Autoregressive foundation models for electronic health records (EHRs) typically inherit pretraining methods from language modeling, where patient trajectories are concatenated into a single token stream and windows are sampled from that stream.

By Joshua Placidi, Yuxuan Liu, Jinpei Han, Marek Rei, A. Aldo Faisal
arXiv AI
Jul 24

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction

arXiv:2607. 20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge, especially for smaller locally deployable models.

By Jessica Sena, Shesadree Priyadarshani, Miguel Contreras, Bharat Gandhi, Scott Siegel, Subhash Nerella, Parisa Rashidi
arXiv Machine Learning
Jun 8

One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models

arXiv:2602. 00541v2 Announce Type: replace Abstract: Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory values or treatment dosages.

By Zilin Jing, Vincent Jeanselme, Yuta Kobayashi, Simon A. Lee, Chao Pang, Aparajita Kashyap, Yanwei Li, Xinzhuo Jiang, Shalmali Joshi
arXiv Computation and Language
Aug 25

Scaling Electronic Health Record Foundation Models for Population Health Management

The paper introduces Scaling Electronic Health Record Foundation Models for Population Health Management, a large‑scale model trained on billions of medical events from over 5 million patients in Taiwan and the United States. By aligning ICD codes across different health systems, the model achieves strong scaling and generalization across 11 chronic disease prediction tasks, outperforming tree‑based, general, and biomedical language models with high sensitivity at 99% specificity. It also demonstrates superior few‑shot performance on the EHRShot benchmark and shows that cross‑system alignment provides a stronger pretraining signal than single‑site duplication in data‑limited scenarios.

By Liwen Sun, Hao-Ren Yao, Ophir Frieder, Xiang Qian, Chenyan Xiong
arXiv Machine Learning
Aug 14

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

arXiv:2608. 12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identification.

By Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto, David Rehkopf, Ayin Vala, Tanmoy Sarkar Pias
arXiv AI
Aug 11

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

arXiv:2505. 16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability.

By Vincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang, Florent Pollet, Young Sang Choi, Xinzhuo Jiang, Yuta Kobayashi, Yanwei Li, Sara Matijevic, Karthik Natarajan, Shalmali Joshi
arXiv AI
Aug 7

Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

arXiv:2608. 06265v1 Announce Type: new Abstract: Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access.

By Omid Bazgir, Md Nasir, Jacob Hoffman, Yang Yang, Manu Agrawal, Anusua Trivedi, Vinay Rao Dandin, Chris Gibbons, Christine Swisher