arXiv Machine Learning

Insulin4RL: Real-Time Insulin Management in the Intensive Care Unit for Offline Reinforcement Learning

arXiv:2606. 19481v1 Announce Type: new Abstract: Offline reinforcement learning (ORL) offers the potential to improve the quality of clinical decision-making using historical electronic health record (EHR) data.

arXiv Machine Learning
Jun 2

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning

arXiv:2606. 01028v1 Announce Type: new Abstract: Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals.

By Yuepeng Wang, Ken Kawano, Yongqi Zhou, Yoshihiko Fujisawa, Richard Weiss, Akifumi Wachi, Katsuki Fujisawa, Ying Chen, Mehrshad Sadria, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
arXiv Machine Learning
Sep 22

A Unified Benchmark for Dynamic Medical Treatment Reinforcement Learning

The paper introduces MedGym, a benchmark environment for dynamic medical treatment recommendation that models patient evolution in continuous time using Physics-Informed Neural Networks. It addresses gaps in existing reinforcement learning (RL) approaches by allowing evaluation of RL methods under irregular measurement intervals, personalized treatment responses, and safety considerations. MedGym enables direct comparison between discrete-time and continuous-time RL methods and supports clinically relevant metrics such as personalization and trajectory-level safety.

By Yuepeng Wang, Ken Kawano, Yoshihiko Fujisawa, Yongqi Zhou, Akifumi Wachi, Mehrshad Sadria, Lei Zhou, Richard Weiss, Katsuki Fujisawa, Ying Chen, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
arXiv Machine Learning
Sep 22

The Evidence Ladder for Reinforcement Learning in Healthcare: From Retrospective Policies to Trusted Interventions

The paper introduces an evidence ladder for evaluating reinforcement learning (RL) in healthcare, outlining stages from problem formulation to lifecycle monitoring. It argues that success in historical data does not guarantee real‑world improvement and highlights assumptions and failure modes at each rung. The authors propose reporting practices to support cumulative evaluation and emphasize that RL should be tested as an intervention within a dynamic sociotechnical system.

By Yunfan Zhao
arXiv AI
3d ago

EHR2Trace: Auditable EHR Data Infrastructure for Patient World Models and Clinical Agents

EHR2Trace is a system that transforms electronic health records from multiple sources into a standardized, traceable event format suitable for training and evaluating patient world models and clinical agents. It links each event to its original record, separates the event time from the time the information became available, and distinguishes between medication orders, dispensing, and administration. The tool supports both OMOP and MEDS data models, includes automated validation, and was tested on three clinical datasets, converting 846.4 million events and detecting all injected faults.

By Xinye Yang, Yuli Wang, Cheng Ting Lin, Harrison Bai
arXiv Machine Learning
Sep 14

Reinforcement Learning over Patient Trajectories for Clinical Reasoning in EHR Foundation Models

The paper proposes a reinforcement learning (RL) fine‑tuning framework for electronic health record (EHR) foundation models, treating them as generative policies over patient trajectories. By framing clinical prediction tasks such as hospital readmission as event‑conditioned, time‑windowed reasoning problems and designing time‑aware, rollout‑sensitive rewards, the authors show that RL fine‑tuning consistently outperforms pre‑trained backbones and strong baselines. The approach enables smaller models to surpass larger pre‑trained models in data‑limited settings, induces positive transfer across tasks, and produces trajectories with stronger structural and semantic alignment to ground truth, improving downstream utility.

By Yuxin Xiao, Sheng Zhang, Chandan Singh, Tristan Naumann, Hoifung Poon, Jianfeng Gao, Xiaodong Liu
arXiv Machine Learning
Sep 22

CNA: An AI-Oriented Comprehensive Normalized Assessment for Healthy Status and Application to Optimize RRT Strategies by Reinforcement Learning

The paper introduces CNA, an AI‑oriented comprehensive normalized assessment that transforms vital‑sign distributions into a standard normal space to produce a unified health‑status score. This score serves as a criterion for evaluating and terminating reinforcement learning policies in Renal Replacement Therapy (RRT). Using a 23‑dimensional state representation and matrix decomposition to handle missing data, the authors apply offline RL algorithms and show that the learned strategy reduces mortality from 13.2% to 5.0% and shortens hospital stays by nearly 59 hours compared to physician‑observed treatments.

By Jiang Liu, Chan Zhou, Yujie Li, Di Wu, Yihao Xie, Peiwei Li, Xin Shu, Jiaqi Zhu, Chunyong Yang, Yuwen Chen, Bin Yi