The paper presents a study on adaptive chemotherapy control using deep reinforcement learning (DRL) to address tumor heterogeneity and drug resistance. Closed‑loop DRL dosing policies—continuous (TD3) and discrete (DQN)—are trained on a high‑dimensional heterogeneous tumor model and benchmarked against a Pontryagin's Maximum Principle (PMP) open‑loop solution. Across a 100‑patient virtual cohort with ±10% parameter perturbations, TD3 achieves higher average tumor reduction, while DQN offers tighter inter‑patient dosing consistency, highlighting an efficacy‑consistency trade‑off. The work assumes full observation of tumor subpopulations, noting that clinical translation will require handling sparse, noisy measurements.
By Bereket Sitotaw Kidane, Md Samiul Haque Motayed, Shuo Wang
arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.
By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
CARE‑VI introduces a framework for improving value targets in off‑policy actor‑critic learning by combining Conservative Adaptive Ranking and Screening (CARS), Selector‑Evaluator Value Assessment (SEVA), and Dynamic Adaptive Risk‑aware Enhancement (DARE). CARS limits candidate actions to a budgeted prefix and expands it only when uncertainty exceeds a threshold; SEVA orders candidates with selector critics and reviews their values with an evaluator critic, capping the value at the selector reference; DARE adjusts residual corrections based on candidate reliability and signal gaps. Theoretical analysis bounds errors in each component, and empirical tests on SAC, TD3, and TD7 across four MuJoCo tasks show CARE‑VI consistently outperforms baselines in mean return.
By Xiang Zou, Shengzhu Shi, Junqi Gao, Zhichang Guo
arXiv:2508. 03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects that continue to shape future states long after the decision that initiated them.
By David Mguni, Wanrong Yang, Jing Dong, Ziquan Liu, Muhammad Salman Haleem, Baoxiang Wang, Dominik Wojtczak
arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.
By Soichiro Nishimori, Paavo Parmas
arXiv:2605. 03065v2 Announce Type: replace Abstract: Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged as effective parameterizations for robot learning.
By Sarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena, Jesse Zhang, Giri Anantharaman, Cleah Winston, Chaoyi Pan, Douglas Chen, Nai-Chieh Huang, Zeynep Temel, Oliver Kroemer, Sergey Levine, Abhishek Gupta, Hongkai Dai, Paarth Shah, Max Simchowitz