arXiv Machine Learning

Adaptive Chemotherapy Control under Tumor Heterogeneity via Reinforcement Learning

The paper presents a study on adaptive chemotherapy control using deep reinforcement learning (DRL) to address tumor heterogeneity and drug resistance. Closed‑loop DRL dosing policies—continuous (TD3) and discrete (DQN)—are trained on a high‑dimensional heterogeneous tumor model and benchmarked against a Pontryagin's Maximum Principle (PMP) open‑loop solution. Across a 100‑patient virtual cohort with ±10% parameter perturbations, TD3 achieves higher average tumor reduction, while DQN offers tighter inter‑patient dosing consistency, highlighting an efficacy‑consistency trade‑off. The work assumes full observation of tumor subpopulations, noting that clinical translation will require handling sparse, noisy measurements.

arXiv Machine Learning
Jul 21

Enhancing Personalized Bladder Cancer Treatment Through Reinforcement Learning: A Recurrent Patient State Transition Decision Support Framework

arXiv:2607. 16916v1 Announce Type: new Abstract: Bladder cancer treatment requires personalized and adaptive decision-making, particularly for recurrent disease, where treatment effectiveness changes across successive clinical episodes.

By Divyansh Chawla, Anshu Garg, Isshaan Singh
arXiv Machine Learning
Jun 2

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning

arXiv:2606. 01028v1 Announce Type: new Abstract: Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals.

By Yuepeng Wang, Ken Kawano, Yongqi Zhou, Yoshihiko Fujisawa, Richard Weiss, Akifumi Wachi, Katsuki Fujisawa, Ying Chen, Mehrshad Sadria, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
arXiv Machine Learning
Sep 22

A Unified Benchmark for Dynamic Medical Treatment Reinforcement Learning

The paper introduces MedGym, a benchmark environment for dynamic medical treatment recommendation that models patient evolution in continuous time using Physics-Informed Neural Networks. It addresses gaps in existing reinforcement learning (RL) approaches by allowing evaluation of RL methods under irregular measurement intervals, personalized treatment responses, and safety considerations. MedGym enables direct comparison between discrete-time and continuous-time RL methods and supports clinically relevant metrics such as personalization and trajectory-level safety.

By Yuepeng Wang, Ken Kawano, Yoshihiko Fujisawa, Yongqi Zhou, Akifumi Wachi, Mehrshad Sadria, Lei Zhou, Richard Weiss, Katsuki Fujisawa, Ying Chen, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen
arXiv Machine Learning
Aug 24

PerturbRx: Learning Treatment-Conditioned Latent Transitions for Patient Drug Response Prediction

PerturbRx is a treatment‑conditioned representation learning framework that learns latent transitions induced by drug interventions. It trains a drug‑ and dose‑conditioned transition predictor using control and treated single‑cell populations, then applies this predictor to pretreatment patient profiles to generate response features without needing post‑treatment data. On TCGA and patient‑derived xenograft benchmarks, PerturbRx outperforms other methods, demonstrating the value of perturbation‑pretrained latent transitions for patient‑level drug‑response prediction.

By Yoshitaka Inoue, Minoh Jeong, Alfred Hero, Rui Kuang, Augustin Luna
arXiv Machine Learning
Sep 4

Multi-step Proximal Policy Improvement in Offline Reinforcement Learning

The paper introduces Multi-step Proximal Policy Improvement (MPI), a method that refines offline reinforcement learning policies through sequential re-centered proximal steps. By viewing policies as a probability manifold, MPI interprets a wide range of offline actor objectives as a single proximal policy improvement step and extends this to multiple steps for controlled policy improvement beyond the behavior distribution. Experiments on D4RL benchmarks demonstrate that a few MPI refinements enhance strong offline baselines such as TD3+BC, ReBRAC, and IQL, while diagnostics clarify the benefits of re-centered refinement over fixed-objective scheduling and highlight critic error limitations.

By Soohyun Choi, Seonvin Cho, Songnam Hong
arXiv Machine Learning
Jun 2

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

arXiv:2606. 01051v1 Announce Type: new Abstract: Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may occur between clinical interactions.

By Xun Shen, Yuepeng Wang, Akifumi Wachi, Yongqi Zhou, Richard Weiss, Yoshihiko Fujisawa, Ken Kawano, Mehrshad Sadria, Ying Chen, Xin Liu, Sebastien Gros, Xiao Hu, Kyoung-Sook Kim, Mengmou Li, Katsuki Fujisawa, Kenji Wakabayashi