arXiv Machine Learning

CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence

arXiv:2608. 03862v1 Announce Type: new Abstract: Emergency triage requires reliable decisions within a short time period.

arXiv AI
6d ago

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

AcuityBench is a new benchmark that tests whether language models can correctly identify the urgency of medical care needed from user presentations. It unifies five public datasets—user conversations, online forum posts, clinical vignettes, and patient portal messages—under a shared four-level acuity framework, providing 914 cases for evaluation. The benchmark supports both explicit four-way classification and free-form conversational responses, revealing that models vary widely in accuracy and that conversational formats reduce over-triage but increase under-triage, especially for high-acuity cases.

By Robin Linzmayer (Department of Computer Science, Columbia University, Department of Biomedical Informatics, Columbia University), Georgianna Lin (Department of Biomedical Informatics, Columbia University), Di Coneybeare (Department of Emergency Medicine, Columbia University Irving Medical Center), Jason Chu (Department of Emergency Medicine, Columbia University Irving Medical Center), Trudi Cloyd (Department of Emergency Medicine, Columbia University Irving Medical Center), Manish Garg (Department of Emergency Medicine, Columbia University Irving Medical Center), Miles Gordon (Department of Emergency Medicine, Columbia University Irving Medical Center), Elizabeth Hartofilis (Department of Emergency Medicine, Columbia University Irving Medical Center), Benjamin Hong (Department of Emergency Medicine, Columbia University Irving Medical Center), Ashraf Hussain (Department of Emergency Medicine, Columbia University Irving Medical Center), Eugene Y. Kim (Department of Emergency Medicine, Columbia University Irving Medical Center), Oluchi Iheagwara King (Department of Emergency Medicine, Columbia University Irving Medical Center), Ross McCormack (Department of Emergency Medicine, Columbia University Irving Medical Center), Erica Olsen (Department of Emergency Medicine, Columbia University Irving Medical Center), John K. Riggins Jr (Department of Emergency Medicine, Columbia University Irving Medical Center), Mustafa N. Rasheed (Department of Emergency Medicine, Columbia University Irving Medical Center), Dana L. Sacco (Department of Emergency Medicine, Columbia University Irving Medical Center), Vinay Saggar (Department of Emergency Medicine, Columbia University Irving Medical Center), Osman R. Sayan (Department of Emergency Medicine, Columbia University Irving Medical Center), Amit Shembekar (Department of Emergency Medicine, Columbia University Irving Medical Center), Janice Shin-Kim (Department of Emergency Medicine, Columbia University Irving Medical Center), Wendy W. Sun (Department of Emergency Medicine, Columbia University Irving Medical Center), Bernard P. Chang (Department of Emergency Medicine, Columbia University Irving Medical Center), David Kessler (Department of Emergency Medicine, Columbia University Irving Medical Center), No\'emie Elhadad (Department of Computer Science, Columbia University, Department of Biomedical Informatics, Columbia University)
arXiv AI
Jul 21

Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare

arXiv:2607. 17508v1 Announce Type: cross Abstract: We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-specific interpretable models that synthesizes coefficient-space structure from natural-language task descriptions and a memory of previously learned task-specific predictors.

By Sazan Mahbub, Caleb Ellington, Zhiyuan Li, Yixin Yang, Souvik Kundu, Ben Lengerich, Eric P. Xing
arXiv AI
Jun 9

TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs

arXiv:2606. 09030v1 Announce Type: cross Abstract: Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk scores for patient triage and interpretable rationales that clinicians can verify.

By Hyeongwon Jang, Gyouk Chu, Changhun Kim, Joonhyung Park, Hangyul Yoon, Eunho Yang
arXiv AI
Aug 11

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

arXiv:2505. 16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability.

By Vincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang, Florent Pollet, Young Sang Choi, Xinzhuo Jiang, Yuta Kobayashi, Yanwei Li, Sara Matijevic, Karthik Natarajan, Shalmali Joshi
arXiv Computation and Language
Sep 11

Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation

The paper introduces the first benchmark for evaluating confidence estimation in large language models during multi‑turn medical consultations, combining three types of medical data and an information sufficiency gradient to capture how confidence and correctness evolve as evidence accumulates. Experiments with 27 methods reveal that token‑level and consistency‑level confidence approaches are limited by medical data, and that medical reasoning must be judged on both diagnostic accuracy and information completeness. Building on these findings, the authors propose MedConf, a retrieval‑augmented, linguistically grounded self‑assessment framework that aligns patient information with supporting, missing, and contradictory relations, producing interpretable confidence estimates that outperform existing methods across multiple datasets and LLMs.

By Zhiyao Ren, Yibing Zhan, Siyuan Liang, Guozheng Ma, Baosheng Yu, Dacheng Tao
arXiv Computation and Language
Sep 22

LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

The study evaluates large language models (LLMs) on sequential emergency department triage, where acuity labels are predicted from progressively longer nurse‑patient conversations. Six LLMs were tested at five checkpoints on simulated and physician‑authored dialogues, showing a decline from moderate‑to‑substantial agreement on full records to only fair‑to‑moderate agreement at each checkpoint. The models consistently anchor on chief complaint exchanges and fail to integrate later evidence, yielding low agreement with clinicians (QWK 0.295 vs. 0.887‑0.929) and concentrating predictions on ESI‑2 and ESI‑3. whyItMatters":"The findings reveal that LLMs, despite strong offline performance, cannot reliably handle the sequential nature of real‑time triage, highlighting a critical gap for safe deployment in emergency settings."

By Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan
arXiv AI
Aug 3

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.

By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar