The paper introduces a bounded updating framework for ICU prediction models that freezes the physiological encoder while allowing updates only to the treatment pathway and fusion head. Experiments on 84,792 MIMIC-IV ICU stays across four temporal shifts show that selective adaptation yields more stable explanations—higher rank correlation, better top‑5 feature agreement, and improved retrieval stability—compared to full model adaptation. Predictive performance varies by outcome, but the results demonstrate that explanation stability is governed by the structural boundaries of allowed adaptation rather than merely by freezing components.
By Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud
arXiv:2510. 15127v3 Announce Type: replace-cross Abstract: Identifying the effects of mechanical ventilation (MV) protocols in critical care requires analyzing data from heterogeneous patient-ventilator systems in the clinical decision-making environment.
By David J. Albers, Tell D. Bennett, Jana de Wiljes, George Hripcsak, Bradford J. Smith, Peter D. Sottile, J. N. Stroh
arXiv:2607. 19020v1 Announce Type: cross Abstract: Background: Clinical decision support systems degrade silently as treatment protocols evolve, yet standard adaptation methods treat models as monolithic blocks, unable to distinguish stable patient physiology from shifting institutional practice.
By Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud
arXiv:2607. 19020v2 Announce Type: replace-cross Abstract: Clinical decision support degrades as treatment protocols evolve, but the obstacle to updating a deployed model is governance as much as accuracy: once retraining touches every parameter, no one can say afterwards where the update acted.
By Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud
The paper reports that counterfactual fairness audits of clinical language‑model agents are unreliable without accounting for a per‑action instability floor. By repeatedly running identical vignettes, the authors found that actions changed 8.7% of the time, with instability varying eightfold across actions. A second model confirmed a pooled floor of 6.7%, showing that any reported fairness estimate lacking this floor cannot be interpreted as evidence of disparity.
By Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar, Rahul Joshi
arXiv:2607. 26752v1 Announce Type: new Abstract: Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning.
By Behraj Khan, Shabir Ahmad, Syed Ahmad Chan Bukhari, Tahir Qasim Syed
arXiv:2609.13543v1 Announce Type: new
Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different...
By Grace Chang Yuan, Xiaoman Zhang, Sung Eun Kim, Luyang Luo, Pranav Rajpurkar
arXiv:2608. 13518v1 Announce Type: new Abstract: Many clinical prediction models treat post-intervention outcomes as a one-step mapping from baseline measurements to a future endpoint.
By Yunsung Chung, Yingshuo Liu, Abboud F. Hassan, Han Feng, Mary M. Maleckar, Nassir Marrouche, Jihun Hamm
The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.
By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
The study evaluates large language models (LLMs) on sequential emergency department triage, where acuity labels are predicted from progressively longer nurse‑patient conversations. Six LLMs were tested at five checkpoints on simulated and physician‑authored dialogues, showing a decline from moderate‑to‑substantial agreement on full records to only fair‑to‑moderate agreement at each checkpoint. The models consistently anchor on chief complaint exchanges and fail to integrate later evidence, yielding low agreement with clinicians (QWK 0.295 vs. 0.887‑0.929) and concentrating predictions on ESI‑2 and ESI‑3.
whyItMatters":"The findings reveal that LLMs, despite strong offline performance, cannot reliably handle the sequential nature of real‑time triage, highlighting a critical gap for safe deployment in emergency settings."
By Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan
arXiv:2606. 05797v1 Announce Type: new Abstract: Longitudinal treatment decisions require predicting potential outcomes under future treatment sequences in the presence of time-varying confounding, heterogeneous patient dynamics, and limited domain-specific data.
By Amirhossein Zare, Amirhessam Zare, Herlock Rahimi, Reza Salarikia, Mohammad Kashkooli
arXiv:2607. 27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science.
By Dennis Thumm, Billy Tim Anthony, Ying Chen