arXiv:2608. 07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy.
By Valentin Li\'{e}vin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang
arXiv:2606. 17474v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly considered for use in clinical consultation tasks, yet most medical evaluations remain static, single-turn, or narrowly outcome-based, limiting their ability to reflect the sequential, uncertain, and interactive nature of real-world care.
By Jiahui Niu, Huizi Yu, Wenkong Wang, Guangxin Dai, Jingxian He, Xiang Li, Zhiying Liang, Xinxin Lin, Kent CY So, Bryan YP Yan, Yun Kwok Wing, Yanqiu Xing, Xin Ma, Lizhou Fan
arXiv:2603. 25821v3 Announce Type: replace-cross Abstract: We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions.
By Anna Kozlova, Stanislau Salavei, Pavel Satalkin, Hanna Plotnitskaya, Sergey Parfenyuk, Andy Nkansah
arXiv:2603. 25821v2 Announce Type: replace-cross Abstract: We present Doctorina MedBench, a comprehensive evaluation framework for agent-based medical AI based on the simulation of realistic physician-patient interactions.
By Anna Kozlova, Stanislau Salavei, Pavel Satalkin, Hanna Plotnitskaya, Sergey Parfenyuk
The study investigates whether coded dialogue logs from generative AI-powered virtual patients can provide teacher-interpretable evidence of clinical reasoning. Analyzing 1,030 dialogues from 210 second-year medical learners, the researchers applied behavioural prevalence, Epistemic Network Analysis, and Transition Network Analysis to identify process patterns linked to high-rated history-taking performance. Findings show that high-rated consultations involve more integrated information gathering, communication, and synthesis, rather than merely increased volume of activity.
By Xinyu Li, Zijian Li, Mengyu Xia, Luzhen Tang, Naping Chen, Changmin Lin, Danijela Gasevic, Dragan Gasevic, Yizhou Fan
The paper introduces MACD, a Multi-Agent Clinical Diagnosis framework that enables large language models to self‑learn clinical knowledge through a multi‑agent pipeline of summarization, refinement, and application. MACD is extended into a human‑AI collaborative workflow where multiple diagnostician agents consult iteratively, guided by a judge agent and human oversight. Evaluation on the MIMIC‑MACD cohort shows significant gains in diagnostic accuracy—an average 11.6 percentage‑point improvement over authoritative knowledge for open‑weight LLMs and an 18.3‑percentage‑point boost over physician‑only diagnosis in text‑only vignettes.
By Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.
By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv:2606. 28900v1 Announce Type: new Abstract: Doctor agents are moving beyond single-turn answer generation toward evolving clinical decision systems.
By Hui Zhang
arXiv:2601. 16529v4 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based guidelines.
By Dongshen Peng, Yi Wang, Austin Schoeffler, Sun-ha Hong, Brian Suffoletto, David Kim, Carl Preiksaitis, Christian Rose
The paper introduces Debate-Mixture-of-Agents (DMoA), a multi‑agent framework that structures role‑based interactions to mimic iterative diagnostic reasoning in clinical settings. Evaluated on 297 rare disease cases and 1,719 challenging cases, DMoA outperformed a GPT‑4o baseline, improving most likely diagnosis accuracy by 10.21 percentage points and safety rate by 11.36 percentage points. Ablation studies and further analyses revealed that these gains stem from the structured workflow rather than merely adding more models or longer outputs, and that performance benefits are influenced by the chosen structure, base model strength, and token budget.
By Chang Xia, Leilei Ouyang, Huimin Wang, Yong Zhao, Kang Li
arXiv:2512. 01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized.
By David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.
By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz