arXiv AI

Agentic AI Enhances Physician Trust in Clinical Decision Making

arXiv:2606. 30658v1 Announce Type: cross Abstract: Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outputs transparent to users.

arXiv AI
Sep 3

Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning

The paper investigates how human interventions at specific fault points—moments when an AI agent’s reasoning is most vulnerable—affect the diagnostic accuracy of multi‑agent medical systems. Using the MedQA dataset, the authors found that correct interventions can boost baseline accuracy by up to 40%, whereas incorrect or bias‑related interventions can reduce performance by up to 6% and increase diagnostic drift and uncertainty. The study also highlights behavioral parallels between cognitive biases observed in simulated agent conversations and real‑world clinical practice, such as premature closure and susceptibility to misleading cues.

By Benjamin C Liu, Dillon Mehta, Rishi Malhotra, Adam Zobian, Yong Ying Tan, Samir Chopra, Daniella Rand, Natalie Pang, Abhiram Gudimella, Kevin Zhu
arXiv AI
Aug 25

The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing

The paper discusses how autonomous AI systems are moving from advisory to agentic roles in medication prescribing, citing recent U.S. legislation and a Utah pilot program. It argues that three architectural features—calibrated per‑prediction confidence, clear differentiation between epistemic and aleatoric uncertainty, and inferential transparency—are essential for safe autonomous prescribing. A survey of 136 U.S. clinicians shows they require a confidence‑based escalation mechanism, prefer different handling of uncertainty types, and will only accept liability when transparency allows informed decision‑making.

By Eileanor LaRocco, Sarah Tan, Adarsh Subbaswamy, Anne Andrews, Andrew Taylor, Cree Gaskin, Chirag Agarwal
arXiv AI
Jul 29

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

arXiv:2607. 25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf.

By Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv AI
Jun 30

An AI agent for treatment reasoning over a biomedical tool universe

arXiv:2606. 28692v1 Announce Type: new Abstract: Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy.

By Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik
arXiv AI
Aug 20

Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering

The paper introduces an adaptive memory and reflection (AMR) multi‑agent system for medical question answering. Each agent has dedicated memory and uses reflection‑based feedback to retrieve relevant prior cases, improving reasoning. The system routes questions through solo, collaborative, or escalated workflows and includes consensus and ethical overseer modules, achieving strong performance on MedQA and MedMCQA datasets.

By Pradeep Murugesan, Luoxiao Yang, Xueli Chen, Xinqi Fan
arXiv AI
Jul 14

Information-seeking failures of large language models in agentic clinical reasoning

arXiv:2607. 10275v1 Announce Type: new Abstract: Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty.

By Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jacqueline Lammert, Sebastian Ziegelmayer, Marcus Makowski, Lisa C. Adams, Keno K. Bressem