arXiv Machine Learning

EHR2Path: Comprehensive Pathway-Level Modeling of Longitudinal Patient Trajectories from Multimodal Electronic Health Records

arXiv:2506. 04831v3 Announce Type: replace Abstract: Forecasting how a patient's condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions, could support more proactive and personalized care, but requires modeling heterogeneous and longitudinal electronic health record (EHR) data.

arXiv Machine Learning
Jun 8

One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models

arXiv:2602. 00541v2 Announce Type: replace Abstract: Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory values or treatment dosages.

By Zilin Jing, Vincent Jeanselme, Yuta Kobayashi, Simon A. Lee, Chao Pang, Aparajita Kashyap, Yanwei Li, Xinzhuo Jiang, Shalmali Joshi
arXiv Machine Learning
Aug 19

Mr.Dec: Daily-Scale Longitudinal Multimodal Modeling for 30-Day Readmission Prediction

Mr.Dec is a new Transformer‑decoder model that predicts 30‑day hospital readmission by treating each admission as a chronological sequence of daily multimodal events, integrating Electronic Health Record updates and Chest X‑ray findings. It uses disease‑specific supervised contrastive learning to shape a diagnosis‑aware latent space and preserves day‑level clinical signals that other methods often compress. Experiments on MIMIC‑IV and MIMIC‑CXR datasets show state‑of‑the‑art performance and the model can highlight "Critical Days" for actionable real‑time risk stratification.

By Minjun Kim, Jong Hak Moon
Hugging Face Trending Papers
Sep 8

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

NOAH is a generative transformer that models the entire multimodal patient journey by integrating bidirectional time and a variational latent space to capture continuous, stochastic clinical trajectories. Trained on over 559 million events from 431,000 hospital visits, it processes medical images, time‑series, numeric signals, categorical events, and both structured and unstructured records. The model supports autoregressive forecasting, zero‑shot classification, and counterfactual simulations, yielding strong predictive performance across 15 ICD chapters, 29 comorbidities, and time‑to‑event outcomes.

arXiv Machine Learning
Aug 21

Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records

arXiv:2608. 20315v1 Announce Type: new Abstract: Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical events.

By Jun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont, Ziv Bar-Joseph, Sven Jager, Brandon Rufino
arXiv Machine Learning
Jul 20

LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

arXiv:2607. 15447v1 Announce Type: new Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods.

By Jingteng Li, Alexander Capstick, Louise Rigny, Iona Biggart, Neil J Sebire, Payam Barnaghi
arXiv Machine Learning
Jun 24

PORTER: Language-Grounded Event Representations for Portable Structured EHR Foundation Models

arXiv:2606. 24102v1 Announce Type: cross Abstract: Most electronic health record (EHR) foundation models encode clinical events as discrete event tokens from a fixed vocabulary and therefore cannot directly represent events containing unseen concepts or new combinations of concepts and attributes such as numeric values.

By Lin Lawrence Guo, Adam Paul Yan, Emily Vettese, Lillian Sung
arXiv AI
Sep 10

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

NOAH is a generative transformer that learns the full multimodal patient journey by integrating bidirectional time and a variational latent space. Trained on over 559 million clinical events from 431,000 hospital visits, it processes medical images, time‑series, numeric signals, categorical events, and both structured and unstructured records. The model supports autoregressive forecasting, zero‑shot classification, and counterfactual intervention simulation, yielding strong performance on clinical outcomes, ICD chapters, comorbidities, and time‑to‑event prediction.

By Tobias Susetzky, Raphael Rehms, Dmitrii Seletkov, \"Ozg\"un Turgut, Michelle Espranita Liman, Lisa Steinhelfer, Rickmer Braren, Daniel Rueckert
arXiv Machine Learning
Aug 19

MultiSigBERT: Beyond Survival Analysis through Multimodal and Sequential Modeling in Oncology

MultiSigBERT is a unified framework that performs multimodal sequential survival modeling in oncology by integrating narrative clinical reports, numerical measurements, and structured variables. The method converts free-text reports into sentence embeddings, compresses them with modality-specific PCA, and concatenates them with structured covariates to create joint temporal trajectories. These trajectories are encoded using the Signature transform from Rough Paths theory, and the resulting high-dimensional features are fed into a LASSO-regularized Cox model, achieving a concordance index of 0.743 on an independent test set of over 2,500 patients.

By Paul Minchella, St\'ephane Chr\'etien, Guillaume Metzler, Lo\"ic Verlingue, R\'emi Vaucher
arXiv AI
Aug 28

Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs

The paper introduces structured evidence routing for incident risk prediction using multimodal longitudinal electronic health records (EHRs). It proposes a router‑predictor‑reviewer workflow that condenses full patient records into compact summaries and targeted evidence slices, enabling a predictor to generate evidence‑linked risk assessments that a reviewer can critique. Experiments on five one‑year incident diagnosis tasks show that the method matches the AUROC of established supervised EHRSHOT baselines and remains competitive on AUPRC, while providing a patient‑specific evidence trail.

By Animesh Agarwal, Meysam Ghaffari, Nina Fatehi, Carlos Morato