The paper introduces ACTMED, a diagnostic framework that combines Bayesian Experimental Design with large language models to emulate real‑world clinical reasoning. ACTMED actively selects the most informative test at each step, using LLMs to simulate patient states and update beliefs without needing task‑specific training data. The authors evaluate the system on real datasets, demonstrating improvements in diagnostic accuracy, interpretability, and efficient resource use while keeping clinicians involved in the decision loop.
By Silas Ruhrberg Est\'evez, Nicol\'as Astorga, Mihaela van der Schaar
arXiv:2602.01995v2 Announce Type: replace
Abstract: Conversational diagnosis requires multi-turn history-taking, where an agent asks clarifying questions to refine differential diagnoses under incomp...
By Jeongmoon Won, Seungwon Kook, Yohan Jo
arXiv:2608.21948v1 Announce Type: new
Abstract: Complex clinical reasoning requires models to update diagnostic hypotheses as new evidence emerges and to coordinate different medical specialities und...
By Sike Xiang, Shuang Chen, Qian sun, Jia Cheng, Yusi Wei, Amir Atapour-Abarghouei
arXiv:2608.22899v1 Announce Type: new
Abstract: Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquire...
By Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Ziru Niu, Zuozhu Liu
arXiv:2603. 14158v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior during clinical interactions.
By Ivan Lopez, Selin S. Everett, Bryan J. Bunning, April S. Liang, Dong Han Yao, Shivam C. Vedak, Kameron C. Black, Sophie Ostmeier, Stephen P. Ma, Emily Alsentzer, Jonathan H. Chen, Akshay S. Chaudhari, Eric Horvitz
arXiv:2509.19375v2 Announce Type: replace-cross
Abstract: Large language models are increasingly used for clinical text classification, where overconfident misclassifications can directly affect pati...
By Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil
arXiv:2606. 08938v1 Announce Type: cross Abstract: Clinical diagnosis requires flexible use of multiple reasoning paradigms under incomplete patient information.
By Gen Li, Yuanze Hu, Zhichao Yang, Qingchen Yu, Jianwei Lv, Yue Guo, Yujing Liu, Faguo Wu, Hongwei Zheng, Xiandong Li, Bo Yuan, Yifan Sun, Zhaoxin Fan
arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.
By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
arXiv:2609.24480v1 Announce Type: cross
Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the conv...
By Kalash Shah, Kunal Singh, Snehan J, Shreyas Singh
arXiv:2608.29582v1 Announce Type: cross
Abstract: Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigati...
By Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi
The paper introduces MACD, a Multi-Agent Clinical Diagnosis framework that enables large language models to self‑learn clinical knowledge through a multi‑agent pipeline of summarization, refinement, and application. MACD is extended into a human‑AI collaborative workflow where multiple diagnostician agents consult iteratively, guided by a judge agent and human oversight. Evaluation on the MIMIC‑MACD cohort shows significant gains in diagnostic accuracy—an average 11.6 percentage‑point improvement over authoritative knowledge for open‑weight LLMs and an 18.3‑percentage‑point boost over physician‑only diagnosis in text‑only vignettes.
By Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
The paper introduces the first benchmark for evaluating confidence estimation in large language models during multi‑turn medical consultations, combining three types of medical data and an information sufficiency gradient to capture how confidence and correctness evolve as evidence accumulates. Experiments with 27 methods reveal that token‑level and consistency‑level confidence approaches are limited by medical data, and that medical reasoning must be judged on both diagnostic accuracy and information completeness. Building on these findings, the authors propose MedConf, a retrieval‑augmented, linguistically grounded self‑assessment framework that aligns patient information with supporting, missing, and contradictory relations, producing interpretable confidence estimates that outperform existing methods across multiple datasets and LLMs.
By Zhiyao Ren, Yibing Zhan, Siyuan Liang, Guozheng Ma, Baosheng Yu, Dacheng Tao