arXiv:2609.27987v1 Announce Type: new
Abstract: Large language models (LLMs) have made substantial progress on medical question-answering, yet effective medical dialogue also requires learning to ask...
By Chenxuan Li, Jiayi Wan, Xinrong Chen, Zhongyu Zhao, Xuecheng Shang, Peixing Wan
arXiv:2607. 18999v1 Announce Type: cross Abstract: Multi-turn medical consultation agents must decide what to ask, adapt to patient responses, and determine when the collected evidence is sufficient.
By Guofeng Zhang, Yizeng Quan, Huaiyi Fang, Jianwei Lv, Jinyao Liu, Xunxu Duan, Lening An, Yu Ouyang, Junfeng Wang
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
By Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng
MIRA is a bilingual benchmark that evaluates whether large language models (LLMs) provide consistent medical information across different user phrasings, languages, and health literacy levels. It contains 4,320 prompts derived from 60 medically reviewed low‑risk health questions and reveals that models tend to omit key information and offer fewer concrete next steps when responding to low health‑literacy signals, a phenomenon termed Differential Information Dilution (DID). A knowledge‑guided mitigation prompt can reduce this dilution for most models, notably improving Claude and Qwen.
By Mengyu Xu, Qiaoxin Yang, Qianqian Wang, Xiwei Dai, Weiyi Wu, Chongyang Gao
arXiv:2608. 19875v1 Announce Type: cross Abstract: Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response.
By Mahyar Abbasian, Saba A. Farahani, Arshia Ilaty, Hung Cao, Ramesh Jain, Amir M. Rahmani
arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.
By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
The study introduces a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) platform designed to train medical students in patient interviews. In a randomized controlled trial with 100 students, the multi-agent system—comprising a patient agent, a Socratic tutor agent, and a turn-level evaluator—did not improve diagnostic accuracy but significantly enhanced overall OSCE scores, especially in communication, empathy, and history-taking. The authors also release a richly annotated dataset to support further research in AI-supported clinical reasoning training.
By Luming Yang, Haoxian Liu, Siqing Li, Rong Jia, Yue Xiao, Guanhua Chen, Li Lu
MedConceal is a new benchmark for evaluating medical dialogue systems on hidden‑concern reasoning under partial observability. It features 300 curated cases and 600 clinician‑LLM interactions, using an interactive patient simulator that hides latent concerns and tracks their revelation and resolution through theory‑grounded communication signals. The benchmark assesses both confirmation (surfacing hidden concerns) and intervention (addressing the primary concern), revealing that current models excel on different metrics while human clinicians still outperform them on intervention success.
By Yikun Han, Joey Chan, Jingyuan Chen, Mengting Ai, Simo Du, Yue Guo
arXiv:2609.09684v1 Announce Type: new
Abstract: Medical question-answering datasets often contain answer labels, whereas high-quality rationales remain scarce, noisy, or costly to validate. This chan...
By Yuexin Wu, Dayou Yu, Vasile Rus
ConsultMind is an uncertainty‑aware framework that automates diagnostic consultation by updating disorder posteriors after each patient response and using posterior uncertainty to guide inquiry and diagnosis. It builds on AutoDisym, a pipeline that constructs a Disorder–Symptom Bayesian Network (DSBN) from diagnostic knowledge and clinical narratives. Across psychiatry, respiratory medicine, fever clinics, and public datasets, AutoDisym produces high‑quality DSBNs and ConsultMind improves diagnostic accuracy and explanation quality, achieving up to 22.15‑point gains in Top‑1 accuracy and 37.89‑point gains in Top‑3 accuracy.
By Xiao Sun, Yuming Yang, Yun Chen, Jiang Zhong, Junnan Zhu, Xinyi Jiang, Haoyang Zeng, Ruirui Chen, Yining Wang, Xinyu Zhou, Rong Tang, Kaiwen Wei
arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.
By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv:2609.24480v1 Announce Type: cross
Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the conv...
By Kalash Shah, Kunal Singh, Snehan J, Shreyas Singh