arXiv:2607. 08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning.
By Fan Ma, Mauro Giuffr\`e, Donald Wright, Kent McCann, Mark Iscoe, Lingfei Qian, Mingyang Jiang, Chi Wing Ng, Na Hong, Huan He, Cathy Shyr, Qingyu Chen, Lee Schwamm, Lucila Ohno-Machado, Hua Xu
arXiv:2607. 15280v1 Announce Type: new Abstract: Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering.
By Shaoting Tan, Ning Liu, Yuntao Du, Shuyue Wei, Wu Shuai, Qian Li, Yanyu Xu, Wei Zhang, Lizhen Cui, Haitao Yuan
arXiv:2608.29582v1 Announce Type: cross
Abstract: Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigati...
By Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi
EviDx is a new framework for evidence-aware active diagnosis that pairs patient-specific diagnostic environments with a clinical scaffold and an observer-guided runtime harness. The framework constructs interactive environments from raw clinical cases, organizes role-specialized agents and evidence tools, and regulates diagnostic termination by tracking uncertainty and evidence coverage. Experiments demonstrate that EviDx improves diagnostic performance and process stability while revealing model-dependent capability boundaries.
By Lihang Zeng, Shaoting Zhang, Xiaofan Zhang
arXiv:2606. 18068v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) and multi-agent systems have driven the rise of Agentic AI, showing promise for medical reasoning.
By Divyansh Srivastava, Shreya Ghosh, Anshul Verma, Rajkumar Buyya
arXiv:2608.21948v1 Announce Type: new
Abstract: Complex clinical reasoning requires models to update diagnostic hypotheses as new evidence emerges and to coordinate different medical specialities und...
By Sike Xiang, Shuang Chen, Qian sun, Jia Cheng, Yusi Wei, Amir Atapour-Abarghouei
VeriDx is a disease‑centric verification framework that links free‑form diagnostic reasoning to structured disease profiles, tracking whether each hypothesis satisfies, remains unresolved, or violates its clinical obligations. It exposes failures such as missing critical tests, unresolved differentials, ignored contradictions, unsupported claims, and premature closure. Applied to complex respiratory diagnosis, VeriDx reveals that many diagnostic errors stem from broken commitments made earlier in the reasoning process.
By Zhong Cao, Shuying Chen
arXiv:2607. 28788v1 Announce Type: new Abstract: Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence.
By Jiahui Li, Ruili Fang, Zishuai Liu, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou
arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.
By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
The paper introduces MIMIC-DOS, a dataset derived from MIMIC-IV that focuses on ICU cases where patient symptoms and medical signs are discordant. It presents CARE, a privacy‑compliant multi‑stage agentic reasoning framework that uses a proprietary LLM to generate structured categories and transitions, while a local LLM performs evidence acquisition and decision‑making. In retrospective evaluations on MIMIC‑DOS, CARE outperforms other LLMs and agentic workflows, demonstrating stronger handling of conflicting clinical evidence while preserving patient privacy.
By Haochen Liu, Weien Li, Rui Song, Zeyu Li, Chun Jason Xue, Xiao-Yang Liu, Sam Nallaperuma-Herzberg, Xue Liu, Ye Yuan
arXiv:2607. 04907v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) in high-stakes clinical settings remains limited by structural hallucinations, weak deterministic reasoning over tabular patient data, and omissions in vector retrieval.
By Mohammed Saim Ahmed Quadri, Yunzhe Xue, Justin W. Ady, Usman Roshan
MTDiag is a newly released multi-turn diagnostic dialogue dataset designed to evaluate large language models (LLMs) in clinically meaningful ways. It is built from DDXPlus, MIMIC-IV, and AJCR case reports, covering both common emergency department presentations and rare conditions, and normalizes cases into a canonical schema using UMLS concept identifiers and ICD-10 codes. The dataset includes a UserLM‑8B utterance‑generation pipeline and physician‑validated natural‑language utterances, and introduces clinical knowledge‑grounded metrics that go beyond simple diagnostic accuracy for multi‑turn differential diagnosis tasks.
By Pia Chouayfati, Alexander M. Fichtl, Miriam Ansch\"utz, George Doumat, Georg Groh