arXiv Computation and Language

Beyond Information Seeking: Severity-Aware Question Supervision for Proactive Medical Dialogue

The paper introduces Expected‑Severity‑Risk (ESR), a new objective for selecting questions in proactive medical dialogue that prioritizes reducing the expected severity of diagnostic errors rather than merely uncertainty. ESR uses population statistics to marginalize over possible answers and distills its rankings into a prefix‑only language policy, enabling deployment without teacher‑side risk computation. Experiments on DDxPlus show ESR cuts high‑severity diagnostic misses by 29.5% and boosts accuracy while adding only 0.14 extra questions per dialogue.

arXiv AI
Sep 4

MIRA: A Bilingual Benchmark for Medical Information Response Audit

MIRA is a bilingual benchmark that evaluates whether large language models (LLMs) provide consistent medical information across different user phrasings, languages, and health literacy levels. It contains 4,320 prompts derived from 60 medically reviewed low‑risk health questions and reveals that models tend to omit key information and offer fewer concrete next steps when responding to low health‑literacy signals, a phenomenon termed Differential Information Dilution (DID). A knowledge‑guided mitigation prompt can reduce this dilution for most models, notably improving Claude and Qwen.

By Mengyu Xu, Qiaoxin Yang, Qianqian Wang, Xiwei Dai, Weiyi Wu, Chongyang Gao
arXiv AI
Aug 3

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.

By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
arXiv AI
Sep 12

Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training

The study introduces a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) platform designed to train medical students in patient interviews. In a randomized controlled trial with 100 students, the multi-agent system—comprising a patient agent, a Socratic tutor agent, and a turn-level evaluator—did not improve diagnostic accuracy but significantly enhanced overall OSCE scores, especially in communication, empathy, and history-taking. The authors also release a richly annotated dataset to support further research in AI-supported clinical reasoning training.

By Luming Yang, Haoxian Liu, Siqing Li, Rong Jia, Yue Xiao, Guanhua Chen, Li Lu
arXiv Computation and Language
Sep 1

MedConceal: A Benchmark for Clinical Hidden-Concern Reasoning Under Partial Observability

MedConceal is a new benchmark for evaluating medical dialogue systems on hidden‑concern reasoning under partial observability. It features 300 curated cases and 600 clinician‑LLM interactions, using an interactive patient simulator that hides latent concerns and tracks their revelation and resolution through theory‑grounded communication signals. The benchmark assesses both confirmation (surfacing hidden concerns) and intervention (addressing the primary concern), revealing that current models excel on different metrics while human clinicians still outperform them on intervention success.

By Yikun Han, Joey Chan, Jingyuan Chen, Mengting Ai, Simo Du, Yue Guo
arXiv AI
3d ago

ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning

ConsultMind is an uncertainty‑aware framework that automates diagnostic consultation by updating disorder posteriors after each patient response and using posterior uncertainty to guide inquiry and diagnosis. It builds on AutoDisym, a pipeline that constructs a Disorder–Symptom Bayesian Network (DSBN) from diagnostic knowledge and clinical narratives. Across psychiatry, respiratory medicine, fever clinics, and public datasets, AutoDisym produces high‑quality DSBNs and ConsultMind improves diagnostic accuracy and explanation quality, achieving up to 22.15‑point gains in Top‑1 accuracy and 37.89‑point gains in Top‑3 accuracy.

By Xiao Sun, Yuming Yang, Yun Chen, Jiang Zhong, Junnan Zhu, Xinyi Jiang, Haoyang Zeng, Ruirui Chen, Yining Wang, Xinyu Zhou, Rong Tang, Kaiwen Wei
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam