arXiv Machine Learning

A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction

arXiv:2608. 04180v1 Announce Type: new Abstract: Feature selection is a critical step in electronic health record (EHR)-based predictive modeling, where input variables are often high-dimensional, sparse, noisy, and redundant.

arXiv Machine Learning
Sep 14

Patient-Reported Survey Data Improve Prediction of Opioid Use Disorder

The study examined whether adding patient‑reported survey data to electronic health records (EHRs) improves the prediction of a first opioid use disorder (OUD) diagnosis. Using 267,747 All of Us participants, the authors compared EHR‑only models to EHR+survey models across multiple machine‑learning algorithms and look‑back windows. Survey augmentation consistently increased predictive performance, with the best 24‑month LightGBM model’s PR‑AUC rising from 0.6219 to 0.6603, and survey features ranked as the second most important information domain.

By Xiyue Jiang, Zihan Ding, Grace Han, Yinan Liu, Richard N. Rosenthal, Fusheng Wang
arXiv AI
Jun 30

Primary ICD Category Prediction using LLM-based Probing

arXiv:2606. 28798v1 Announce Type: new Abstract: Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables.

By Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen
arXiv Machine Learning
Aug 4

xMICD: Explainable Representation of Multiple ICD Codes

arXiv:2608. 00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning.

By Pat Vatiwutipong, Kumkup Keeratisiwakul, Albert Phuoc Kien Van Truong, Nutcha Yodrabum, Wasin Pansiritanachot, Marvin N. Wright, Thanapon Noraset
arXiv Machine Learning
Sep 17

Enhancing Extubation Failure Prediction with LLM-Derived Features from Respiratory Therapy Clinical Notes

The paper proposes a new method for predicting extubation failure (EF) by extracting features from free-text respiratory therapy notes using a large language model and combining them with logistic regression. Applied to a cohort from University of Washington Medicine, the approach identifies clinically relevant EF-related features that enhance prediction performance when added to structured patient data. The study also discusses how varying target populations and EF definitions in prior research can cause systematic performance differences and limit generalizability.

By Izzy Chaiken, Aditya Khowal, Neha A. Sathe, Mark M. Wurfel, Lucy Lu Wang
arXiv Machine Learning
Aug 4

EHR2Path: Comprehensive Pathway-Level Modeling of Longitudinal Patient Trajectories from Multimodal Electronic Health Records

arXiv:2506. 04831v3 Announce Type: replace Abstract: Forecasting how a patient's condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions, could support more proactive and personalized care, but requires modeling heterogeneous and longitudinal electronic health record (EHR) data.

By Chantal Pellegrini, Ege \"Ozsoy, David Bani-Harouni, Matthias Keicher, Nassir Navab
arXiv Machine Learning
Jun 8

One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models

arXiv:2602. 00541v2 Announce Type: replace Abstract: Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory values or treatment dosages.

By Zilin Jing, Vincent Jeanselme, Yuta Kobayashi, Simon A. Lee, Chao Pang, Aparajita Kashyap, Yanwei Li, Xinzhuo Jiang, Shalmali Joshi
arXiv Machine Learning
Aug 31

Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining

The paper refactors and expands the scikit-rebate Python package, adding new Relief‑Based Algorithm (RBA) variants such as SWRF*, mu‑Relief, and five novel methods that use alternative neighbor selection and feature scoring strategies. Benchmarking across diverse genomic simulations shows that most RBAs, except mu‑Relief, effectively detect 2‑way interactions in noisy data, with far‑scoring variants like MultiSWRFDB* excelling at interaction detection but being less sensitive to main effects. The refactored package achieves 10‑ to 35‑fold runtime reductions, and the new RBAs maintain strong performance for both main effects and 2‑way epistatic interactions, preserving predictive signals for downstream modeling.

By Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz