arXiv:2602. 10385v5 Announce Type: replace-cross Abstract: The contemporary paradigm of trajectory learning operates fundamentally at the level of group dynamics, systematically reducing individual-level complexity to fit group-level models, thus rendering effective patient subtyping difficult and individual-level modeling largely out of reach.
By Jia Li, Yu Hou, Rui Zhang
Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors.
arXiv:2608. 12779v1 Announce Type: cross Abstract: Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment.
By Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang
arXiv:2605. 07267v2 Announce Type: replace Abstract: Personalized healthcare decisions require reasoning about how physiological and behavioral variables influence an individual patient over time.
By Elahe Khatibi, Ziyu Wang, Saba A. Farahani, Di Huang, Hung Cao, Ramesh Jain, Amir M. Rahmani
arXiv:2608. 10339v1 Announce Type: cross Abstract: Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions.
By Patrick Vossler, Jialin Ouyang, F. Richard Guo, Anran Huang, Ali Shojaie, Lucas Zier, Fan Xia, Jean Feng
arXiv:2606. 29503v1 Announce Type: cross Abstract: The verbose context problem occurs when structured concepts have token-inefficient textual representations.
By Shiva Kaul, Min-Gyu Kim, Anjum Khurshid, Sriram Vishwanath
The paper proposes a reinforcement learning (RL) fine‑tuning framework for electronic health record (EHR) foundation models, treating them as generative policies over patient trajectories. By framing clinical prediction tasks such as hospital readmission as event‑conditioned, time‑windowed reasoning problems and designing time‑aware, rollout‑sensitive rewards, the authors show that RL fine‑tuning consistently outperforms pre‑trained backbones and strong baselines. The approach enables smaller models to surpass larger pre‑trained models in data‑limited settings, induces positive transfer across tasks, and produces trajectories with stronger structural and semantic alignment to ground truth, improving downstream utility.
By Yuxin Xiao, Sheng Zhang, Chandan Singh, Tristan Naumann, Hoifung Poon, Jianfeng Gao, Xiaodong Liu
The paper introduces a paired benchmark to detect hindsight bias in clinical language models by comparing model responses to questions posed at a clinically relevant cutoff versus the full timeline. It uses 171 case reports (40 sepsis, 131 GLP‑1/diabetes) with both human‑annotated and LLM‑generated time‑series data, evaluating accuracy, hindsight trap rate, answer instability rate, and hindsight bias rate. Results show that exposing models to the full timeline consistently increases hindsight bias, while truncating the timeline mitigates bias without sacrificing accuracy.
By Misaki Matsuura, Sayantan Kumar, Ojas Kadam, Jeremy C. Weiss
The paper introduces information set emulation, a method that attaches detailed causal certificates—such as source evidence, timing, and proposed causal roles—to AI‑derived features extracted from electronic health records (EHRs). These certificates provide auditable evidence for causal roles and guide whether a feature can be used for causal inference or should be routed to compatible reporting or separate analyses. The framework integrates with a joint EHR observation map and offers identification, estimation, and diagnostic tools under standard causal assumptions, illustrated through synthetic simulations and a finite‑world example.
By Takes Fujita (VRI), Nobutaka Hattori (Department of Neurology, Juntendo University School of Medicine)
The paper "Medical Causal Hypothesis Verification with Large Language Models" reports a small-scale study evaluating eight LLMs on 17 medical causal hypotheses. The authors introduce an evaluation framework and annotate 1,067 evidence points across six criteria, using nine metrics to assess performance. Results show that while LLMs have strong recall, they frequently fail to provide valid scientific articles, evidence, or reject unsupported hypotheses, revealing a critical limitation for their use in healthcare.
By Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam, Elena Zheleva
Structural causal models are the standard language for reasoning about interventions and counterfactuals, but they describe static variables, typically measured once, and usually forbid cyclic dependencies. Many systems we care about, such as patients, climates, and economies, instead evolve continuously in time, are observed at irregular time points, and contain feedback loops.
The verbose context problem occurs when structured concepts have token-inefficient textual representations. This bottleneck is acute in population health: cohort-level analysis of longitudinal patient records requires reasoning over thousands of medically-coded events, often exceeding 400K tokens in total.