arXiv:2606. 17405v1 Announce Type: new Abstract: Clinical decision support AI systems (CDSASs) must adapt to evolving patient conditions in real-time while adhering to strict safety constraints.
By Xinyu Qin, Anil K. Sood, Ruiheng Yu, Sara Corvigno, Elaine Stur, Lu Wang
The paper presents a study on adaptive chemotherapy control using deep reinforcement learning (DRL) to address tumor heterogeneity and drug resistance. Closed‑loop DRL dosing policies—continuous (TD3) and discrete (DQN)—are trained on a high‑dimensional heterogeneous tumor model and benchmarked against a Pontryagin's Maximum Principle (PMP) open‑loop solution. Across a 100‑patient virtual cohort with ±10% parameter perturbations, TD3 achieves higher average tumor reduction, while DQN offers tighter inter‑patient dosing consistency, highlighting an efficacy‑consistency trade‑off. The work assumes full observation of tumor subpopulations, noting that clinical translation will require handling sparse, noisy measurements.
By Bereket Sitotaw Kidane, Md Samiul Haque Motayed, Shuo Wang
arXiv:2512.08029v4 Announce Type: replace
Abstract: Clinical decision-making in oncology requires forecasting how disease evolves under treatment, yet most AI systems remain static predictors that ca...
By Tianxingjian Ding, Yuanhao Zou, Chen Chen, Mubarak Shah, Yu Tian
arXiv:2607. 16916v1 Announce Type: new Abstract: Bladder cancer treatment requires personalized and adaptive decision-making, particularly for recurrent disease, where treatment effectiveness changes across successive clinical episodes.
By Divyansh Chawla, Anshu Garg, Isshaan Singh
arXiv:2606. 01051v1 Announce Type: new Abstract: Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may occur between clinical interactions.
By Xun Shen, Yuepeng Wang, Akifumi Wachi, Yongqi Zhou, Richard Weiss, Yoshihiko Fujisawa, Ken Kawano, Mehrshad Sadria, Ying Chen, Xin Liu, Sebastien Gros, Xiao Hu, Kyoung-Sook Kim, Mengmou Li, Katsuki Fujisawa, Kenji Wakabayashi
arXiv:2606. 01028v1 Announce Type: new Abstract: Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals.
By Yuepeng Wang, Ken Kawano, Yongqi Zhou, Yoshihiko Fujisawa, Richard Weiss, Akifumi Wachi, Katsuki Fujisawa, Ying Chen, Mehrshad Sadria, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
The paper introduces MedGym, a benchmark environment for dynamic medical treatment recommendation that models patient evolution in continuous time using Physics-Informed Neural Networks. It addresses gaps in existing reinforcement learning (RL) approaches by allowing evaluation of RL methods under irregular measurement intervals, personalized treatment responses, and safety considerations. MedGym enables direct comparison between discrete-time and continuous-time RL methods and supports clinically relevant metrics such as personalization and trajectory-level safety.
By Yuepeng Wang, Ken Kawano, Yoshihiko Fujisawa, Yongqi Zhou, Akifumi Wachi, Mehrshad Sadria, Lei Zhou, Richard Weiss, Katsuki Fujisawa, Ying Chen, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
arXiv:2512. 08029v3 Announce Type: replace Abstract: Clinical decision-making in oncology requires predicting dynamic disease evolution, a task current static AI predictors cannot perform.
By Tianxingjian Ding, Yuanhao Zou, Chen Chen, Mubarak Shah, Yu Tian
arXiv:2606. 25762v1 Announce Type: new Abstract: In oncology, access to patient-level data is often restricted.
By Octavia-Andreea Ciora, Julian Welzel, Dennis Frauen, Maresa Schr\"oder, Marie Brockschmidt, Harry Amad, Thomas Callender, Mihaela van der Schaar, Stefan Feuerriegel
PerturbRx is a treatment‑conditioned representation learning framework that learns latent transitions induced by drug interventions. It trains a drug‑ and dose‑conditioned transition predictor using control and treated single‑cell populations, then applies this predictor to pretreatment patient profiles to generate response features without needing post‑treatment data. On TCGA and patient‑derived xenograft benchmarks, PerturbRx outperforms other methods, demonstrating the value of perturbation‑pretrained latent transitions for patient‑level drug‑response prediction.
By Yoshitaka Inoue, Minoh Jeong, Alfred Hero, Rui Kuang, Augustin Luna
arXiv:2608. 03606v1 Announce Type: new Abstract: Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence.
By William Bolton, Philip Torr
arXiv:2607. 08793v1 Announce Type: cross Abstract: Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested.
By Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung