Hugging Face Trending Papers

Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation

Read the original on Hugging Face Trending Papers →

The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned policy cannot be trialed on patients, its value must be estimated off-policy, and such estimates can be fragile and optimistic.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 18

Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation

arXiv:2608. 16482v1 Announce Type: new Abstract: The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care.

By Marc P\'erez-Roig, David Fern\'andez-Narro, Carlos S\'aez
arXiv Machine Learning
Aug 13

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

arXiv:2608. 11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which agents replicate harmful patterns such as treatment withdrawal during comfort-care transitions.

By Hangqi Ren, Junyi Liao
arXiv AI
Aug 28

Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study

The study presents a new sepsis severity score derived from 43 routinely charted variables over a 72‑hour window, trained on 29,116 and 7,691 adult patients from two Massachusetts hospitals. Using mortality as a treatment‑level ranking signal, the index assigns higher scores to non‑survivors (1.19–1.64 points higher on a 0–10 scale) across various baseline strata and shows significant correlations with lactate, MAP, and creatinine changes. Cross‑institutional agreement and external validation demonstrate robust performance, suggesting the score could serve as a decision‑support tool alongside clinical judgment.

By Kevin Zhu, Ryan Zhang, Baraa Abed, Tilendra Choudhary, Malvern Madondo, Mehak Arora, Yixuan Yang, Alasdair Gent, Aditya Nagori, Omer T. Inan, Krista L. Haines, Patrick Georgoff, Suresh M. Agarwal, Vijay Krishnamoorthy, Tetsu Ohnuma, Mihai V. Podgoreanu, Michael R. Pinsky, Gilles Clermont, Craig M. Coopersmith, Craig S. Jabaley, Rishikesan Kamaleswaran
arXiv Machine Learning
1d ago

Patient-Centered Treatment Planning for Chronic Multimorbidity: A Hierarchical Reinforcement Learning Framework for Preference Modeling

The paper introduces FAHOC, a hierarchical reinforcement learning framework that models patient preferences by learning high‑level therapeutic options and factored intra‑option policies, while enforcing a cooperation‑aware action masking mechanism. It demonstrates that cooperative patients achieve better health outcomes and that the Q‑function approximation error is bounded. Evaluated on data from ~50,000 comorbid hypertension and type 2 diabetes patients, FAHOC improves quality‑adjusted life years by 0.669, correctly identifies cooperative patients 95.9% of the time, and never violates patient preferences in held‑out tests.

By Nafiseh Payani, Soham Das, G. Anthony Wilson, Anahita Khojandi
arXiv Machine Learning
Aug 4

Inferring Relative Consequences of Mechanical Ventilation from Observational Data Using Game-Based Comparisons

arXiv:2510. 15127v3 Announce Type: replace-cross Abstract: Identifying the effects of mechanical ventilation (MV) protocols in critical care requires analyzing data from heterogeneous patient-ventilator systems in the clinical decision-making environment.

By David J. Albers, Tell D. Bennett, Jana de Wiljes, George Hripcsak, Bradford J. Smith, Peter D. Sottile, J. N. Stroh