arXiv:2608. 11420v1 Announce Type: new Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users.
By Del Coburn, Scott Sanner, Dan Silver
The paper introduces MACD, a Multi-Agent Clinical Diagnosis framework that enables large language models to self‑learn clinical knowledge through a multi‑agent pipeline of summarization, refinement, and application. MACD is extended into a human‑AI collaborative workflow where multiple diagnostician agents consult iteratively, guided by a judge agent and human oversight. Evaluation on the MIMIC‑MACD cohort shows significant gains in diagnostic accuracy—an average 11.6 percentage‑point improvement over authoritative knowledge for open‑weight LLMs and an 18.3‑percentage‑point boost over physician‑only diagnosis in text‑only vignettes.
By Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
The article surveys LLM-based agentic reasoning frameworks, presenting a unified formal language that categorizes them into single-agent, tool-based, and multi-agent methods. It reviews application scenarios in scientific discovery, healthcare, software engineering, society, economics, and general-purpose tasks, and compares the distinct features and evaluation strategies of each category. The survey highlights the rapid development of complex agentic systems in real-world contexts.
By Bingxi Zhao, Lin Geng Foo, Ping Hu, Christian Theobalt, Hossein Rahmani, Jun Liu
arXiv:2609.15161v1 Announce Type: cross
Abstract: Large language model (LLM) driven multi-agent systems have shown promise in complex clinical reasoning, yet existing approaches rely on static strate...
By Dongsheng Shi, Yue Li, Xin Yi, Linlin Wang
arXiv:2606. 01094v1 Announce Type: new Abstract: Clinical order generation serves as a critical bridge between clinical decision-making and real-world practice, translating medical decisions into concrete and executable orders.
By Ruihui Hou, Ziyue Huai, Chennuo Zhang, Ziyan Liu, Siran Zhao, Yao Yu, Jie Zhai, Tong Ruan
The paper reviews how Large Language Models (LLMs) are being adapted for medical reasoning, moving beyond single-step answers to systems that can systematically, transparently, and verifiably reason. It introduces a taxonomy of enhancement techniques, split into training-time methods such as supervised fine‑tuning and reinforcement learning, and test-time methods like prompt engineering and multi‑agent systems. The review examines their application across text, image, and code modalities in key clinical areas—diagnosis, education, and treatment planning—and tracks the shift in evaluation benchmarks from simple accuracy to more nuanced assessments of reasoning quality and visual interpretability.
By Zizhan Ma, Wenxuan Wang, Meidan Ding, Shiyi Zheng, Shengyuan Liu, Jie Liu, Jiaming Ji, Linlin Shen, Yixuan Yuan, Wenting Chen