The paper introduces MedGym, a benchmark environment for dynamic medical treatment recommendation that models patient evolution in continuous time using Physics-Informed Neural Networks. It addresses gaps in existing reinforcement learning (RL) approaches by allowing evaluation of RL methods under irregular measurement intervals, personalized treatment responses, and safety considerations. MedGym enables direct comparison between discrete-time and continuous-time RL methods and supports clinically relevant metrics such as personalization and trajectory-level safety.
By Yuepeng Wang, Ken Kawano, Yoshihiko Fujisawa, Yongqi Zhou, Akifumi Wachi, Mehrshad Sadria, Lei Zhou, Richard Weiss, Katsuki Fujisawa, Ying Chen, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
arXiv:2606. 01028v1 Announce Type: new Abstract: Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals.
By Yuepeng Wang, Ken Kawano, Yongqi Zhou, Yoshihiko Fujisawa, Richard Weiss, Akifumi Wachi, Katsuki Fujisawa, Ying Chen, Mehrshad Sadria, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
arXiv:2607. 05620v1 Announce Type: cross Abstract: In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold.
By Katherine Avery, Bruno Castro da Silva, David Jensen
arXiv:2607. 08793v1 Announce Type: cross Abstract: Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested.
By Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung
The paper introduces Exchange Policy Optimization (EPO), a framework for semi‑infinite safe reinforcement learning that handles infinitely many constraints by iteratively solving finite subproblems. EPO expands or deletes constraints based on tolerance violations and Lagrange multipliers, maintaining computational tractability while converging to an optimal policy with bounded safety violations. The authors prove finite convergence, provide iteration bounds, and quantify the suboptimality gap under mild assumptions.
By Jiaming Zhang, Yujie Yang, Haoning Wang, Liping Zhang, Shengbo Eben Li
arXiv:2606. 10376v1 Announce Type: new Abstract: Cancer treatment is at the core a sequential decision-making problem with partial observability, latent patient heterogeneity, and explicit constraints on the budget for medical measurements.
By Deniz Sargun, H. Bugra Tulay, C. Emre Koksal
The paper presents a study on adaptive chemotherapy control using deep reinforcement learning (DRL) to address tumor heterogeneity and drug resistance. Closed‑loop DRL dosing policies—continuous (TD3) and discrete (DQN)—are trained on a high‑dimensional heterogeneous tumor model and benchmarked against a Pontryagin's Maximum Principle (PMP) open‑loop solution. Across a 100‑patient virtual cohort with ±10% parameter perturbations, TD3 achieves higher average tumor reduction, while DQN offers tighter inter‑patient dosing consistency, highlighting an efficacy‑consistency trade‑off. The work assumes full observation of tumor subpopulations, noting that clinical translation will require handling sparse, noisy measurements.
By Bereket Sitotaw Kidane, Md Samiul Haque Motayed, Shuo Wang
arXiv:2508. 03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects that continue to shape future states long after the decision that initiated them.
By David Mguni, Wanrong Yang, Jing Dong, Ziquan Liu, Muhammad Salman Haleem, Baoxiang Wang, Dominik Wojtczak
arXiv:2606. 19481v1 Announce Type: new Abstract: Offline reinforcement learning (ORL) offers the potential to improve the quality of clinical decision-making using historical electronic health record (EHR) data.
By Thomas Frost, Steve Harris
The paper proposes a reinforcement learning (RL) fine‑tuning framework for electronic health record (EHR) foundation models, treating them as generative policies over patient trajectories. By framing clinical prediction tasks such as hospital readmission as event‑conditioned, time‑windowed reasoning problems and designing time‑aware, rollout‑sensitive rewards, the authors show that RL fine‑tuning consistently outperforms pre‑trained backbones and strong baselines. The approach enables smaller models to surpass larger pre‑trained models in data‑limited settings, induces positive transfer across tasks, and produces trajectories with stronger structural and semantic alignment to ground truth, improving downstream utility.
By Yuxin Xiao, Sheng Zhang, Chandan Singh, Tristan Naumann, Hoifung Poon, Jianfeng Gao, Xiaodong Liu
GPAgentBench-2K is a new benchmark that evaluates large language model agents in primary‑care clinical decision‑making using a constrained Markov Decision Process (CMDP). It is built from expert‑validated records of real GP encounters and includes six foundational clinical actions, a workflow prior, and safety‑informed abstention as an outcome. Testing 16 state‑of‑the‑art LLMs shows that performance drops as the action space grows and that even the best models violate safety constraints in more than half of high‑risk cases, highlighting a gap between clinical quality and safety.
By Boqi Chen, Xudong Liu, Yunke Ao, Heejin Do, Jianing Qiu
arXiv:2606. 17405v1 Announce Type: new Abstract: Clinical decision support AI systems (CDSASs) must adapt to evolving patient conditions in real-time while adhering to strict safety constraints.
By Xinyu Qin, Anil K. Sood, Ruiheng Yu, Sara Corvigno, Elaine Stur, Lu Wang