The paper introduces MACD, a Multi-Agent Clinical Diagnosis framework that enables large language models to self‑learn clinical knowledge through a multi‑agent pipeline of summarization, refinement, and application. MACD is extended into a human‑AI collaborative workflow where multiple diagnostician agents consult iteratively, guided by a judge agent and human oversight. Evaluation on the MIMIC‑MACD cohort shows significant gains in diagnostic accuracy—an average 11.6 percentage‑point improvement over authoritative knowledge for open‑weight LLMs and an 18.3‑percentage‑point boost over physician‑only diagnosis in text‑only vignettes.
By Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
arXiv:2608. 11420v1 Announce Type: new Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users.
By Del Coburn, Scott Sanner, Dan Silver
arXiv:2608.21948v1 Announce Type: new
Abstract: Complex clinical reasoning requires models to update diagnostic hypotheses as new evidence emerges and to coordinate different medical specialities und...
By Sike Xiang, Shuang Chen, Qian sun, Jia Cheng, Yusi Wei, Amir Atapour-Abarghouei
The paper introduces Debate-Mixture-of-Agents (DMoA), a multi‑agent framework that structures role‑based interactions to mimic iterative diagnostic reasoning in clinical settings. Evaluated on 297 rare disease cases and 1,719 challenging cases, DMoA outperformed a GPT‑4o baseline, improving most likely diagnosis accuracy by 10.21 percentage points and safety rate by 11.36 percentage points. Ablation studies and further analyses revealed that these gains stem from the structured workflow rather than merely adding more models or longer outputs, and that performance benefits are influenced by the chosen structure, base model strength, and token budget.
By Chang Xia, Leilei Ouyang, Huimin Wang, Yong Zhao, Kang Li
arXiv:2607. 15280v1 Announce Type: new Abstract: Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering.
By Shaoting Tan, Ning Liu, Yuntao Du, Shuyue Wei, Wu Shuai, Qian Li, Yanyu Xu, Wei Zhang, Lizhen Cui, Haitao Yuan
arXiv:2609.15161v1 Announce Type: cross
Abstract: Large language model (LLM) driven multi-agent systems have shown promise in complex clinical reasoning, yet existing approaches rely on static strate...
By Dongsheng Shi, Yue Li, Xin Yi, Linlin Wang
arXiv:2608.22899v1 Announce Type: new
Abstract: Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquire...
By Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Ziru Niu, Zuozhu Liu
arXiv:2510. 21324v2 Announce Type: replace Abstract: Chest X-ray (CXR) plays a pivotal role in clinical diagnosis, and a variety of task-specific and foundation models have been developed for automatic CXR interpretation.
By Jinhui Lou, Yan Yang, Zhou Yu, Zhenqi Fu, Weidong Han, Qingming Huang, Jun Yu
EviDx is a new framework for evidence-aware active diagnosis that pairs patient-specific diagnostic environments with a clinical scaffold and an observer-guided runtime harness. The framework constructs interactive environments from raw clinical cases, organizes role-specialized agents and evidence tools, and regulates diagnostic termination by tracking uncertainty and evidence coverage. Experiments demonstrate that EviDx improves diagnostic performance and process stability while revealing model-dependent capability boundaries.
By Lihang Zeng, Shaoting Zhang, Xiaofan Zhang
arXiv:2606. 08093v1 Announce Type: new Abstract: Pathology is the cornerstone of modern medicine, where accurate decision-making relies heavily on evidence-based practices.
By Zhe Xu, Zhengyu Zhang, Zhiyuan Cai, Jiahao Xu, Yijie Lin, Ziyi Liu, Junlin Hou, Hongyi Wang, Yuxiang Nie, Ling Liang, Yihui Wang, Yingxue Xu, Ronald Cheong Kin Chan, Li Liang, Hao Chen
PathPocket is a multimodal AI co‑pilot that grounds pathology decision‑making in evidence. It builds the largest pathology evidence corpus (≈110,472 documents) and a hypergraph of 4.55 million entities and 7.10 million relations to support traceable reasoning. The system handles text and multimodal queries, including ROI and gigapixel whole‑slide images, and outperforms current state‑of‑the‑art models on a benchmark of over 200,000 real‑world cases, improving pathologists’ diagnostic accuracy and confidence.
By Zhe Xu, Zhengyu Zhang, Zhiyuan Cai, Jiahao Xu, Yijie Lin, Ziyi Liu, Junlin Hou, Hongyi Wang, Yuxiang Nie, Yihui Wang, Jiabo Ma, Ling Liang, Yingxue Xu, Zhengrui Guo, Guanghao Wu, Danyi Li, Ziqi Zhou, Donglin Tan, Zhijian Cen, Ying Tan, Xiaolin Liu, Qi Xie, Xiaoying Tang, Xi Peng, Cheng Deng, Lijuan Qu, Ronald Cheong Kin Chan, Li Liang, Hao Chen
arXiv:2606. 08938v1 Announce Type: cross Abstract: Clinical diagnosis requires flexible use of multiple reasoning paradigms under incomplete patient information.
By Gen Li, Yuanze Hu, Zhichao Yang, Qingchen Yu, Jianwei Lv, Yue Guo, Yujing Liu, Faguo Wu, Hongwei Zheng, Xiandong Li, Bo Yuan, Yifan Sun, Zhaoxin Fan