The paper investigates how human interventions at specific fault points—moments when an AI agent’s reasoning is most vulnerable—affect the diagnostic accuracy of multi‑agent medical systems. Using the MedQA dataset, the authors found that correct interventions can boost baseline accuracy by up to 40%, whereas incorrect or bias‑related interventions can reduce performance by up to 6% and increase diagnostic drift and uncertainty. The study also highlights behavioral parallels between cognitive biases observed in simulated agent conversations and real‑world clinical practice, such as premature closure and susceptibility to misleading cues.
By Benjamin C Liu, Dillon Mehta, Rishi Malhotra, Adam Zobian, Yong Ying Tan, Samir Chopra, Daniella Rand, Natalie Pang, Abhiram Gudimella, Kevin Zhu
The paper discusses how autonomous AI systems are moving from advisory to agentic roles in medication prescribing, citing recent U.S. legislation and a Utah pilot program. It argues that three architectural features—calibrated per‑prediction confidence, clear differentiation between epistemic and aleatoric uncertainty, and inferential transparency—are essential for safe autonomous prescribing. A survey of 136 U.S. clinicians shows they require a confidence‑based escalation mechanism, prefer different handling of uncertainty types, and will only accept liability when transparency allows informed decision‑making.
By Eileanor LaRocco, Sarah Tan, Adarsh Subbaswamy, Anne Andrews, Andrew Taylor, Cree Gaskin, Chirag Agarwal
arXiv:2608.30676v1 Announce Type: new
Abstract: When medical AI systems hallucinate clinical reasoning, the consequences extend beyond incorrect answers: fabricated justifications that superficially...
By Jiangwang Chen, Chenghao Zhang, Hengxing Cai
arXiv:2608.29453v1 Announce Type: cross
Abstract: As AI becomes increasingly integrated into clinical practice, it is playing a growing role in medical decision making. Medicine, however, is a high s...
By Jiayuan Zhu, Jiazhen Pan, Fenglin Liu, Minhao Hu, Junde Wu
arXiv:2607. 25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf.
By Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv:2608.21864v1 Announce Type: cross
Abstract: The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these frameworks often...
By Md Asaduzzaman Jabin, Zihao Wu, Tianming Liu
arXiv:2606. 01094v1 Announce Type: new Abstract: Clinical order generation serves as a critical bridge between clinical decision-making and real-world practice, translating medical decisions into concrete and executable orders.
By Ruihui Hou, Ziyue Huai, Chennuo Zhang, Ziyan Liu, Siran Zhao, Yao Yu, Jie Zhai, Tong Ruan
arXiv:2606. 14149v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in healthcare settings, yet their tendency to hallucinate poses risks when clinical decisions are involved.
By Muhammad Osama, Maheera Amjad, Zartasha Mustansar, Arslan Shaukat, Muhammad U. S. Khan
arXiv:2606. 28692v1 Announce Type: new Abstract: Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy.
By Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik
The paper introduces an adaptive memory and reflection (AMR) multi‑agent system for medical question answering. Each agent has dedicated memory and uses reflection‑based feedback to retrieve relevant prior cases, improving reasoning. The system routes questions through solo, collaborative, or escalated workflows and includes consensus and ethical overseer modules, achieving strong performance on MedQA and MedMCQA datasets.
By Pradeep Murugesan, Luoxiao Yang, Xueli Chen, Xinqi Fan
arXiv:2607. 10275v1 Announce Type: new Abstract: Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty.
By Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jacqueline Lammert, Sebastian Ziegelmayer, Marcus Makowski, Lisa C. Adams, Keno K. Bressem
arXiv:2608.23397v1 Announce Type: new
Abstract: Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final diagnosis a...
By Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao