arXiv Computation and Language By Chenxuan Li, Xinrong Chen, Luyan Zhang, Peidong Jia, Runfan Zheng, Zhongyu Zhao, Xuecheng Shang, Peixing Wan

Beyond Information Seeking: Severity-Aware Question Supervision for Proactive Medical Dialogue

Read the original on arXiv Computation and Language →

The paper introduces Expected‑Severity‑Risk (ESR), a new objective for selecting questions in proactive medical dialogue that prioritizes reducing the expected severity of diagnostic errors rather than merely uncertainty. ESR uses population statistics to marginalize over possible answers and distills its rankings into a prefix‑only language policy, enabling deployment without teacher‑side risk computation. Experiments on DDxPlus show ESR cuts high‑severity diagnostic misses by 29.5% and boosts accuracy while adding only 0.14 extra questions per dialogue.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 4

MIRA: A Bilingual Benchmark for Medical Information Response Audit

MIRA is a bilingual benchmark that evaluates whether large language models (LLMs) provide consistent medical information across different user phrasings, languages, and health literacy levels. It contains 4,320 prompts derived from 60 medically reviewed low‑risk health questions and reveals that models tend to omit key information and offer fewer concrete next steps when responding to low health‑literacy signals, a phenomenon termed Differential Information Dilution (DID). A knowledge‑guided mitigation prompt can reduce this dilution for most models, notably improving Claude and Qwen.

By Mengyu Xu, Qiaoxin Yang, Qianqian Wang, Xiwei Dai, Weiyi Wu, Chongyang Gao
arXiv AI
Aug 3

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.

By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar