arXiv AI By Robin Linzmayer (Department of Computer Science, Columbia University, Department of Biomedical Informatics, Columbia University), Georgianna Lin (Department of Biomedical Informatics, Columbia University), Di Coneybeare (Department of Emergency Medicine, Columbia University Irving Medical Center), Jason Chu (Department of Emergency Medicine, Columbia University Irving Medical Center), Trudi Cloyd (Department of Emergency Medicine, Columbia University Irving Medical Center), Manish Garg (Department of Emergency Medicine, Columbia University Irving Medical Center), Miles Gordon (Department of Emergency Medicine, Columbia University Irving Medical Center), Elizabeth Hartofilis (Department of Emergency Medicine, Columbia University Irving Medical Center), Benjamin Hong (Department of Emergency Medicine, Columbia University Irving Medical Center), Ashraf Hussain (Department of Emergency Medicine, Columbia University Irving Medical Center), Eugene Y. Kim (Department of Emergency Medicine, Columbia University Irving Medical Center), Oluchi Iheagwara King (Department of Emergency Medicine, Columbia University Irving Medical Center), Ross McCormack (Department of Emergency Medicine, Columbia University Irving Medical Center), Erica Olsen (Department of Emergency Medicine, Columbia University Irving Medical Center), John K. Riggins Jr (Department of Emergency Medicine, Columbia University Irving Medical Center), Mustafa N. Rasheed (Department of Emergency Medicine, Columbia University Irving Medical Center), Dana L. Sacco (Department of Emergency Medicine, Columbia University Irving Medical Center), Vinay Saggar (Department of Emergency Medicine, Columbia University Irving Medical Center), Osman R. Sayan (Department of Emergency Medicine, Columbia University Irving Medical Center), Amit Shembekar (Department of Emergency Medicine, Columbia University Irving Medical Center), Janice Shin-Kim (Department of Emergency Medicine, Columbia University Irving Medical Center), Wendy W. Sun (Department of Emergency Medicine, Columbia University Irving Medical Center), Bernard P. Chang (Department of Emergency Medicine, Columbia University Irving Medical Center), David Kessler (Department of Emergency Medicine, Columbia University Irving Medical Center), No\'emie Elhadad (Department of Computer Science, Columbia University, Department of Biomedical Informatics, Columbia University)

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

Read the original on arXiv AI →

AcuityBench is a new benchmark that tests whether language models can correctly identify the urgency of medical care needed from user presentations. It unifies five public datasets—user conversations, online forum posts, clinical vignettes, and patient portal messages—under a shared four-level acuity framework, providing 914 cases for evaluation. The benchmark supports both explicit four-way classification and free-form conversational responses, revealing that models vary widely in accuracy and that conversational formats reduce over-triage but increase under-triage, especially for high-acuity cases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 11

Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation

The paper introduces the first benchmark for evaluating confidence estimation in large language models during multi‑turn medical consultations, combining three types of medical data and an information sufficiency gradient to capture how confidence and correctness evolve as evidence accumulates. Experiments with 27 methods reveal that token‑level and consistency‑level confidence approaches are limited by medical data, and that medical reasoning must be judged on both diagnostic accuracy and information completeness. Building on these findings, the authors propose MedConf, a retrieval‑augmented, linguistically grounded self‑assessment framework that aligns patient information with supporting, missing, and contradictory relations, producing interpretable confidence estimates that outperform existing methods across multiple datasets and LLMs.

By Zhiyao Ren, Yibing Zhan, Siyuan Liang, Guozheng Ma, Baosheng Yu, Dacheng Tao
arXiv AI
Jul 21

Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare

arXiv:2607. 17508v1 Announce Type: cross Abstract: We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-specific interpretable models that synthesizes coefficient-space structure from natural-language task descriptions and a memory of previously learned task-specific predictors.

By Sazan Mahbub, Caleb Ellington, Zhiyuan Li, Yixin Yang, Souvik Kundu, Ben Lengerich, Eric P. Xing
arXiv AI
Jul 29

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

arXiv:2607. 25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf.

By Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf