arXiv Machine Learning

Patient-Reported Survey Data Improve Prediction of Opioid Use Disorder

The study examined whether adding patient‑reported survey data to electronic health records (EHRs) improves the prediction of a first opioid use disorder (OUD) diagnosis. Using 267,747 All of Us participants, the authors compared EHR‑only models to EHR+survey models across multiple machine‑learning algorithms and look‑back windows. Survey augmentation consistently increased predictive performance, with the best 24‑month LightGBM model’s PR‑AUC rising from 0.6219 to 0.6603, and survey features ranked as the second most important information domain.

arXiv Machine Learning
Aug 6

A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction

arXiv:2608. 04180v1 Announce Type: new Abstract: Feature selection is a critical step in electronic health record (EHR)-based predictive modeling, where input variables are often high-dimensional, sparse, noisy, and redundant.

By Zihan Ding, Yinan Liu, Tengfei Ma, Rachel Wong, George Leibowitz, Benjamin Littenberg, Xia Zheng, Richard N. Rosenthal, Fusheng Wang
arXiv Machine Learning
Jun 8

One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models

arXiv:2602. 00541v2 Announce Type: replace Abstract: Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory values or treatment dosages.

By Zilin Jing, Vincent Jeanselme, Yuta Kobayashi, Simon A. Lee, Chao Pang, Aparajita Kashyap, Yanwei Li, Xinzhuo Jiang, Shalmali Joshi
arXiv AI
3d ago

EHR2Trace: Auditable EHR Data Infrastructure for Patient World Models and Clinical Agents

EHR2Trace is a system that transforms electronic health records from multiple sources into a standardized, traceable event format suitable for training and evaluating patient world models and clinical agents. It links each event to its original record, separates the event time from the time the information became available, and distinguishes between medication orders, dispensing, and administration. The tool supports both OMOP and MEDS data models, includes automated validation, and was tested on three clinical datasets, converting 846.4 million events and detecting all injected faults.

By Xinye Yang, Yuli Wang, Cheng Ting Lin, Harrison Bai
arXiv AI
Aug 18

Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic

arXiv:2608. 16273v1 Announce Type: cross Abstract: Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research.

By Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson
arXiv AI
Aug 28

Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs

The paper introduces structured evidence routing for incident risk prediction using multimodal longitudinal electronic health records (EHRs). It proposes a router‑predictor‑reviewer workflow that condenses full patient records into compact summaries and targeted evidence slices, enabling a predictor to generate evidence‑linked risk assessments that a reviewer can critique. Experiments on five one‑year incident diagnosis tasks show that the method matches the AUROC of established supervised EHRSHOT baselines and remains competitive on AUPRC, while providing a patient‑specific evidence trail.

By Animesh Agarwal, Meysam Ghaffari, Nina Fatehi, Carlos Morato