The study examined whether adding patient‑reported survey data to electronic health records (EHRs) improves the prediction of a first opioid use disorder (OUD) diagnosis. Using 267,747 All of Us participants, the authors compared EHR‑only models to EHR+survey models across multiple machine‑learning algorithms and look‑back windows. Survey augmentation consistently increased predictive performance, with the best 24‑month LightGBM model’s PR‑AUC rising from 0.6219 to 0.6603, and survey features ranked as the second most important information domain.
By Xiyue Jiang, Zihan Ding, Grace Han, Yinan Liu, Richard N. Rosenthal, Fusheng Wang
arXiv:2603. 02221v2 Announce Type: replace-cross Abstract: In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods.
By Zizheng Zhang, Yiming Li, Justin Xu, Jinyu Wang, Rui Wang, Lei Song, Jiang Bian, David W Eyre, Jingjing Fu
arXiv:2609.15713v1 Announce Type: new
Abstract: Recent approaches to 30-day hospital readmission prediction rely on pre-trained language models applied to discharge summaries. Although these methods...
By Mohamad Najafi, Hongyun Fu, Mathias Brochhausen, Jian Wu, Yaohang Li
arXiv:2603.02221v3 Announce Type: replace-cross
Abstract: In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods. LLMs are increasi...
By Zizheng Zhang, Yiming Li, Justin Xu, Jinyu Wang, Rui Wang, Lei Song, Jiang Bian, David W Eyre, Jingjing Fu
arXiv:2609.40108v1 Announce Type: new
Abstract: Opioid overdose remains a major clinical and public health burden, highlighting the need for scalable approaches to identify patients at high risk. Her...
By Mingchen Li, Rohan Pandey, Junhui Qian, Feiyun Ouyang, Sunjae Kwon, Avijit Mitra, Zonghai Yao, Hong Yu
arXiv:2606. 28798v1 Announce Type: new Abstract: Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables.
By Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen