The study examined whether adding patient‑reported survey data to electronic health records (EHRs) improves the prediction of a first opioid use disorder (OUD) diagnosis. Using 267,747 All of Us participants, the authors compared EHR‑only models to EHR+survey models across multiple machine‑learning algorithms and look‑back windows. Survey augmentation consistently increased predictive performance, with the best 24‑month LightGBM model’s PR‑AUC rising from 0.6219 to 0.6603, and survey features ranked as the second most important information domain.
By Xiyue Jiang, Zihan Ding, Grace Han, Yinan Liu, Richard N. Rosenthal, Fusheng Wang
arXiv:2603. 02221v2 Announce Type: replace-cross Abstract: In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods.
By Zizheng Zhang, Yiming Li, Justin Xu, Jinyu Wang, Rui Wang, Lei Song, Jiang Bian, David W Eyre, Jingjing Fu
arXiv:2609.15713v1 Announce Type: new
Abstract: Recent approaches to 30-day hospital readmission prediction rely on pre-trained language models applied to discharge summaries. Although these methods...
By Mohamad Najafi, Hongyun Fu, Mathias Brochhausen, Jian Wu, Yaohang Li
arXiv:2603.02221v3 Announce Type: replace-cross
Abstract: In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods. LLMs are increasi...
By Zizheng Zhang, Yiming Li, Justin Xu, Jinyu Wang, Rui Wang, Lei Song, Jiang Bian, David W Eyre, Jingjing Fu
arXiv:2609.40108v1 Announce Type: new
Abstract: Opioid overdose remains a major clinical and public health burden, highlighting the need for scalable approaches to identify patients at high risk. Her...
By Mingchen Li, Rohan Pandey, Junhui Qian, Feiyun Ouyang, Sunjae Kwon, Avijit Mitra, Zonghai Yao, Hong Yu
arXiv:2606. 28798v1 Announce Type: new Abstract: Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables.
By Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen
arXiv:2608. 00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning.
By Pat Vatiwutipong, Kumkup Keeratisiwakul, Albert Phuoc Kien Van Truong, Nutcha Yodrabum, Wasin Pansiritanachot, Marvin N. Wright, Thanapon Noraset
The paper proposes a new method for predicting extubation failure (EF) by extracting features from free-text respiratory therapy notes using a large language model and combining them with logistic regression. Applied to a cohort from University of Washington Medicine, the approach identifies clinically relevant EF-related features that enhance prediction performance when added to structured patient data. The study also discusses how varying target populations and EF definitions in prior research can cause systematic performance differences and limit generalizability.
By Izzy Chaiken, Aditya Khowal, Neha A. Sathe, Mark M. Wurfel, Lucy Lu Wang
arXiv:2506. 04831v3 Announce Type: replace Abstract: Forecasting how a patient's condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions, could support more proactive and personalized care, but requires modeling heterogeneous and longitudinal electronic health record (EHR) data.
By Chantal Pellegrini, Ege \"Ozsoy, David Bani-Harouni, Matthias Keicher, Nassir Navab
arXiv:2602. 00541v2 Announce Type: replace Abstract: Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory values or treatment dosages.
By Zilin Jing, Vincent Jeanselme, Yuta Kobayashi, Simon A. Lee, Chao Pang, Aparajita Kashyap, Yanwei Li, Xinzhuo Jiang, Shalmali Joshi
arXiv:2607. 09165v1 Announce Type: cross Abstract: Achieving early and timely diagnosis and treatment for disease is a major challenge.
By Qingchu Jin, Felistas Mazhude, Jamie B. Rabb, Robert S. Kramer, Douglas B. Sawyer, Raimond L. Winslow
The paper refactors and expands the scikit-rebate Python package, adding new Relief‑Based Algorithm (RBA) variants such as SWRF*, mu‑Relief, and five novel methods that use alternative neighbor selection and feature scoring strategies. Benchmarking across diverse genomic simulations shows that most RBAs, except mu‑Relief, effectively detect 2‑way interactions in noisy data, with far‑scoring variants like MultiSWRFDB* excelling at interaction detection but being less sensitive to main effects. The refactored package achieves 10‑ to 35‑fold runtime reductions, and the new RBAs maintain strong performance for both main effects and 2‑way epistatic interactions, preserving predictive signals for downstream modeling.
By Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz