arXiv AI

ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning

ConsultMind is an uncertainty‑aware framework that automates diagnostic consultation by updating disorder posteriors after each patient response and using posterior uncertainty to guide inquiry and diagnosis. It builds on AutoDisym, a pipeline that constructs a Disorder–Symptom Bayesian Network (DSBN) from diagnostic knowledge and clinical narratives. Across psychiatry, respiratory medicine, fever clinics, and public datasets, AutoDisym produces high‑quality DSBNs and ConsultMind improves diagnostic accuracy and explanation quality, achieving up to 22.15‑point gains in Top‑1 accuracy and 37.89‑point gains in Top‑3 accuracy.

arXiv AI
Sep 12

Timely Clinical Diagnosis through Active Test Selection

The paper introduces ACTMED, a diagnostic framework that combines Bayesian Experimental Design with large language models to emulate real‑world clinical reasoning. ACTMED actively selects the most informative test at each step, using LLMs to simulate patient states and update beliefs without needing task‑specific training data. The authors evaluate the system on real datasets, demonstrating improvements in diagnostic accuracy, interpretability, and efficient resource use while keeping clinicians involved in the decision loop.

By Silas Ruhrberg Est\'evez, Nicol\'as Astorga, Mihaela van der Schaar
arXiv Machine Learning
Aug 7

Clinician input steers AI toward accurate and harmful recommendations

arXiv:2603. 14158v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior during clinical interactions.

By Ivan Lopez, Selin S. Everett, Bryan J. Bunning, April S. Liang, Dong Han Yao, Shivam C. Vedak, Kameron C. Black, Sophie Ostmeier, Stephen P. Ma, Emily Alsentzer, Jonathan H. Chen, Akshay S. Chaudhari, Eric Horvitz
arXiv AI
Aug 18

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.

By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
arXiv AI
Aug 25

MACD: Multi-Agent Clinical Diagnosis with Self-Learned Knowledge for LLM

The paper introduces MACD, a Multi-Agent Clinical Diagnosis framework that enables large language models to self‑learn clinical knowledge through a multi‑agent pipeline of summarization, refinement, and application. MACD is extended into a human‑AI collaborative workflow where multiple diagnostician agents consult iteratively, guided by a judge agent and human oversight. Evaluation on the MIMIC‑MACD cohort shows significant gains in diagnostic accuracy—an average 11.6 percentage‑point improvement over authoritative knowledge for open‑weight LLMs and an 18.3‑percentage‑point boost over physician‑only diagnosis in text‑only vignettes.

By Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
arXiv Computation and Language
Sep 11

Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation

The paper introduces the first benchmark for evaluating confidence estimation in large language models during multi‑turn medical consultations, combining three types of medical data and an information sufficiency gradient to capture how confidence and correctness evolve as evidence accumulates. Experiments with 27 methods reveal that token‑level and consistency‑level confidence approaches are limited by medical data, and that medical reasoning must be judged on both diagnostic accuracy and information completeness. Building on these findings, the authors propose MedConf, a retrieval‑augmented, linguistically grounded self‑assessment framework that aligns patient information with supporting, missing, and contradictory relations, producing interpretable confidence estimates that outperform existing methods across multiple datasets and LLMs.

By Zhiyao Ren, Yibing Zhan, Siyuan Liang, Guozheng Ma, Baosheng Yu, Dacheng Tao