arXiv:2608. 16482v1 Announce Type: new Abstract: The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care.
By Marc P\'erez-Roig, David Fern\'andez-Narro, Carlos S\'aez
arXiv:2608. 11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which agents replicate harmful patterns such as treatment withdrawal during comfort-care transitions.
By Hangqi Ren, Junyi Liao
The study presents a new sepsis severity score derived from 43 routinely charted variables over a 72‑hour window, trained on 29,116 and 7,691 adult patients from two Massachusetts hospitals. Using mortality as a treatment‑level ranking signal, the index assigns higher scores to non‑survivors (1.19–1.64 points higher on a 0–10 scale) across various baseline strata and shows significant correlations with lactate, MAP, and creatinine changes. Cross‑institutional agreement and external validation demonstrate robust performance, suggesting the score could serve as a decision‑support tool alongside clinical judgment.
By Kevin Zhu, Ryan Zhang, Baraa Abed, Tilendra Choudhary, Malvern Madondo, Mehak Arora, Yixuan Yang, Alasdair Gent, Aditya Nagori, Omer T. Inan, Krista L. Haines, Patrick Georgoff, Suresh M. Agarwal, Vijay Krishnamoorthy, Tetsu Ohnuma, Mihai V. Podgoreanu, Michael R. Pinsky, Gilles Clermont, Craig M. Coopersmith, Craig S. Jabaley, Rishikesan Kamaleswaran
arXiv:2607. 08793v1 Announce Type: cross Abstract: Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested.
By Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung
The paper introduces FAHOC, a hierarchical reinforcement learning framework that models patient preferences by learning high‑level therapeutic options and factored intra‑option policies, while enforcing a cooperation‑aware action masking mechanism. It demonstrates that cooperative patients achieve better health outcomes and that the Q‑function approximation error is bounded. Evaluated on data from ~50,000 comorbid hypertension and type 2 diabetes patients, FAHOC improves quality‑adjusted life years by 0.669, correctly identifies cooperative patients 95.9% of the time, and never violates patient preferences in held‑out tests.
By Nafiseh Payani, Soham Das, G. Anthony Wilson, Anahita Khojandi
arXiv:2510. 15127v3 Announce Type: replace-cross Abstract: Identifying the effects of mechanical ventilation (MV) protocols in critical care requires analyzing data from heterogeneous patient-ventilator systems in the clinical decision-making environment.
By David J. Albers, Tell D. Bennett, Jana de Wiljes, George Hripcsak, Bradford J. Smith, Peter D. Sottile, J. N. Stroh
arXiv:2606. 04632v1 Announce Type: new Abstract: Mechanical ventilation for Acute Respiratory Distress Syndrome (ARDS) requires balancing competing physiological goals, including oxygenation, lung protection, and acid-base homeostasis.
By Teqi Hao, Yuxuan Fu, Xiaoyu Tan, Shaojie Shi, Bohao Lv, Yinghui Xu, Xihe Qiu
arXiv:2607. 19020v1 Announce Type: cross Abstract: Background: Clinical decision support systems degrade silently as treatment protocols evolve, yet standard adaptation methods treat models as monolithic blocks, unable to distinguish stable patient physiology from shifting institutional practice.
By Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud
arXiv:2608. 07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy.
By Valentin Li\'{e}vin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang
arXiv:2606. 20206v1 Announce Type: cross Abstract: In offline Reinforcement Learning, immediate rewards in logged batch data are often unobserved due to sparse or irregular record-keeping, or censored beyond certain reward values.
By Ziheng Wei, Annie Qu, Rui Miao
arXiv:2608.30442v1 Announce Type: new
Abstract: Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, p...
By Kihun Rhee
Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL al...