arXiv Machine Learning

Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models

The paper investigates how the granularity of interventions—single components versus full treatment bundles—affects counterfactual simulations using a clinical world model. By analyzing 945,707 patient-hours from MIMIC-IV, the authors show that interventions are naturally bundled, and editing a single component often represents an unreal scenario. Experiments with Clin‑JEPA on 1,019 ventilation onsets demonstrate that editing the complete bundle produces larger predicted state changes than editing individual settings, indicating that bundle-aware editing better captures treatment sensitivity.

arXiv AI
Sep 15

Freezing the Physiological Encoder: Explanation Stability Under Bounded Updates of an ICU Model

The paper introduces a bounded updating framework for ICU prediction models that freezes the physiological encoder while allowing updates only to the treatment pathway and fusion head. Experiments on 84,792 MIMIC-IV ICU stays across four temporal shifts show that selective adaptation yields more stable explanations—higher rank correlation, better top‑5 feature agreement, and improved retrieval stability—compared to full model adaptation. Predictive performance varies by outcome, but the results demonstrate that explanation stability is governed by the structural boundaries of allowed adaptation rather than merely by freezing components.

By Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud
arXiv Machine Learning
Aug 4

Inferring Relative Consequences of Mechanical Ventilation from Observational Data Using Game-Based Comparisons

arXiv:2510. 15127v3 Announce Type: replace-cross Abstract: Identifying the effects of mechanical ventilation (MV) protocols in critical care requires analyzing data from heterogeneous patient-ventilator systems in the clinical decision-making environment.

By David J. Albers, Tell D. Bennett, Jana de Wiljes, George Hripcsak, Bradford J. Smith, Peter D. Sottile, J. N. Stroh
arXiv AI
Jul 22

Biological Amnesia in ICU Time-Series Prediction: A Drift-Adaptive Two-Stream Architecture with Temporal Retrieval

arXiv:2607. 19020v1 Announce Type: cross Abstract: Background: Clinical decision support systems degrade silently as treatment protocols evolve, yet standard adaptation methods treat models as monolithic blocks, unable to distinguish stable patient physiology from shifting institutional practice.

By Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud
arXiv AI
Aug 21

Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating

arXiv:2607. 19020v2 Announce Type: replace-cross Abstract: Clinical decision support degrades as treatment protocols evolve, but the obstacle to updating a deployed model is governance as much as accuracy: once retraining touches every parameter, no one can say afterwards where the update acted.

By Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud
arXiv Machine Learning
Sep 4

Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor

The paper reports that counterfactual fairness audits of clinical language‑model agents are unreliable without accounting for a per‑action instability floor. By repeatedly running identical vignettes, the authors found that actions changed 8.7% of the time, with instability varying eightfold across actions. A second model confirmed a pooled floor of 6.7%, showing that any reported fairness estimate lacking this floor cannot be interpreted as evidence of disparity.

By Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar, Rahul Joshi
arXiv Machine Learning
Jul 30

CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation

arXiv:2607. 26752v1 Announce Type: new Abstract: Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning.

By Behraj Khan, Shabir Ahmad, Syed Ahmad Chan Bukhari, Tahir Qasim Syed
arXiv AI
2d ago

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.

By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
arXiv Computation and Language
Sep 22

LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

The study evaluates large language models (LLMs) on sequential emergency department triage, where acuity labels are predicted from progressively longer nurse‑patient conversations. Six LLMs were tested at five checkpoints on simulated and physician‑authored dialogues, showing a decline from moderate‑to‑substantial agreement on full records to only fair‑to‑moderate agreement at each checkpoint. The models consistently anchor on chief complaint exchanges and fail to integrate later evidence, yielding low agreement with clinicians (QWK 0.295 vs. 0.887‑0.929) and concentrating predictions on ESI‑2 and ESI‑3. whyItMatters":"The findings reveal that LLMs, despite strong offline performance, cannot reliably handle the sequential nature of real‑time triage, highlighting a critical gap for safe deployment in emergency settings."

By Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan
arXiv Machine Learning
Jul 31

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

arXiv:2607. 27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science.

By Dennis Thumm, Billy Tim Anthony, Ying Chen