arXiv:2604.26766v2 Announce Type: replace-cross
Abstract: Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly varia...
By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
The study evaluates large language models (LLMs) on sequential emergency department triage, where acuity labels are predicted from progressively longer nurse‑patient conversations. Six LLMs were tested at five checkpoints on simulated and physician‑authored dialogues, showing a decline from moderate‑to‑substantial agreement on full records to only fair‑to‑moderate agreement at each checkpoint. The models consistently anchor on chief complaint exchanges and fail to integrate later evidence, yielding low agreement with clinicians (QWK 0.295 vs. 0.887‑0.929) and concentrating predictions on ESI‑2 and ESI‑3.
whyItMatters":"The findings reveal that LLMs, despite strong offline performance, cannot reliably handle the sequential nature of real‑time triage, highlighting a critical gap for safe deployment in emergency settings."
By Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan
arXiv:2512.19735v4 Announce Type: replace
Abstract: Accurately predicting mortality risk in intensive care unit (ICU) patients is critical for clinical decision-making. Large language models (LLMs) a...
By Gangxiong Zhang, Yongchao Long, Yuxi Zhou, Yong Zhang, Shenda Hong
The study evaluates whether large language models (LLMs) with in‑context learning can better identify institution‑specific protected health information (PHI) in electronic health records than existing de‑identification systems. Using 100 pediatric oncology notes from Texas Children’s Hospital, eight LLMs were compared to two purpose‑built systems and pattern‑based baselines under three progressively specific prompts. The best LLM achieved an F1 score of 0.918, recovering 79% of previously missed PHI categories and reaching a recall of 0.981 after iterative prompt refinement, demonstrating that calibrated single‑pass prompting can close the institutional PHI gap while balancing precision and recall.
By Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong
The paper reports that counterfactual fairness audits of clinical language‑model agents are unreliable without accounting for a per‑action instability floor. By repeatedly running identical vignettes, the authors found that actions changed 8.7% of the time, with instability varying eightfold across actions. A second model confirmed a pooled floor of 6.7%, showing that any reported fairness estimate lacking this floor cannot be interpreted as evidence of disparity.
By Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar, Rahul Joshi
arXiv:2605.01048v2 Announce Type: replace-cross
Abstract: Counterfactual prompting (i.e., perturbing a single factor and measuring output change) is widely used to evaluate things like LLM bias and C...
By Zihao Yang, Mosh Levy, Yoav Goldberg, Byron C. Wallace
arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.
By Koyar Afrasyab
arXiv:2606. 16723v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.
By Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor
arXiv:2607. 13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving.
By Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler
arXiv:2607. 18086v1 Announce Type: new Abstract: Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved.
By Koyar Afrasyab
arXiv:2608. 10273v1 Announce Type: cross Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally deployable open-source small language models (SLMs).
By Qingfeng Zhang, Yuanxiong Guo, Yanmin Gong
arXiv:2608. 03028v1 Announce Type: new Abstract: Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision.
By Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang