arXiv:2606. 01028v1 Announce Type: new Abstract: Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals.
By Yuepeng Wang, Ken Kawano, Yongqi Zhou, Yoshihiko Fujisawa, Richard Weiss, Akifumi Wachi, Katsuki Fujisawa, Ying Chen, Mehrshad Sadria, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
arXiv:2606. 01051v1 Announce Type: new Abstract: Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may occur between clinical interactions.
By Xun Shen, Yuepeng Wang, Akifumi Wachi, Yongqi Zhou, Richard Weiss, Yoshihiko Fujisawa, Ken Kawano, Mehrshad Sadria, Ying Chen, Xin Liu, Sebastien Gros, Xiao Hu, Kyoung-Sook Kim, Mengmou Li, Katsuki Fujisawa, Kenji Wakabayashi
arXiv:2606. 19481v1 Announce Type: new Abstract: Offline reinforcement learning (ORL) offers the potential to improve the quality of clinical decision-making using historical electronic health record (EHR) data.
By Thomas Frost, Steve Harris
arXiv:2607. 16916v1 Announce Type: new Abstract: Bladder cancer treatment requires personalized and adaptive decision-making, particularly for recurrent disease, where treatment effectiveness changes across successive clinical episodes.
By Divyansh Chawla, Anshu Garg, Isshaan Singh
arXiv:2512. 08029v3 Announce Type: replace Abstract: Clinical decision-making in oncology requires predicting dynamic disease evolution, a task current static AI predictors cannot perform.
By Tianxingjian Ding, Yuanhao Zou, Chen Chen, Mubarak Shah, Yu Tian
arXiv:2606. 17405v1 Announce Type: new Abstract: Clinical decision support AI systems (CDSASs) must adapt to evolving patient conditions in real-time while adhering to strict safety constraints.
By Xinyu Qin, Anil K. Sood, Ruiheng Yu, Sara Corvigno, Elaine Stur, Lu Wang
arXiv:2607. 08793v1 Announce Type: cross Abstract: Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested.
By Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung
The paper proposes a reinforcement learning (RL) fine‑tuning framework for electronic health record (EHR) foundation models, treating them as generative policies over patient trajectories. By framing clinical prediction tasks such as hospital readmission as event‑conditioned, time‑windowed reasoning problems and designing time‑aware, rollout‑sensitive rewards, the authors show that RL fine‑tuning consistently outperforms pre‑trained backbones and strong baselines. The approach enables smaller models to surpass larger pre‑trained models in data‑limited settings, induces positive transfer across tasks, and produces trajectories with stronger structural and semantic alignment to ground truth, improving downstream utility.
By Yuxin Xiao, Sheng Zhang, Chandan Singh, Tristan Naumann, Hoifung Poon, Jianfeng Gao, Xiaodong Liu
arXiv:2609.36144v1 Announce Type: new
Abstract: Electronic health records provide irregular observations of latent patient states that evolve continuously over time. Recent autoregressive models cond...
By Silas Ruhrberg Est\'evez, Kara Liu, Christopher Chiu, Benjamin Atta Owusu, Umesh Kadam, Russ B. Altman, Mihaela van der Schaar
arXiv:2606. 16721v1 Announce Type: new Abstract: Medical diagnosis and treatment are dynamic processes in which patient states evolve over time and clinical interventions alter future outcomes.
By Ke Liu, Mengxuan Li, Yanyi Bao, Tianyun Zhang, Chong Chu, Jiajun Bu, Haishuai Wang
The paper presents a study on adaptive chemotherapy control using deep reinforcement learning (DRL) to address tumor heterogeneity and drug resistance. Closed‑loop DRL dosing policies—continuous (TD3) and discrete (DQN)—are trained on a high‑dimensional heterogeneous tumor model and benchmarked against a Pontryagin's Maximum Principle (PMP) open‑loop solution. Across a 100‑patient virtual cohort with ±10% parameter perturbations, TD3 achieves higher average tumor reduction, while DQN offers tighter inter‑patient dosing consistency, highlighting an efficacy‑consistency trade‑off. The work assumes full observation of tumor subpopulations, noting that clinical translation will require handling sparse, noisy measurements.
By Bereket Sitotaw Kidane, Md Samiul Haque Motayed, Shuo Wang
arXiv:2512.08029v4 Announce Type: replace
Abstract: Clinical decision-making in oncology requires forecasting how disease evolves under treatment, yet most AI systems remain static predictors that ca...
By Tianxingjian Ding, Yuanhao Zou, Chen Chen, Mubarak Shah, Yu Tian