arXiv:2608.23397v1 Announce Type: new
Abstract: Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final diagnosis a...
By Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao
arXiv:2609.13543v1 Announce Type: new
Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different...
By Grace Chang Yuan, Xiaoman Zhang, Sung Eun Kim, Luyang Luo, Pranav Rajpurkar
arXiv:2605. 12895v2 Announce Type: replace-cross Abstract: Clinical decision-support systems are expert systems whose recommendations clinicians act on directly, yet they are usually cleared on one aggregate accuracy number from a held-out test set.
By Rohith Reddy Bellibatlu, Manpreet Singh, Yash Jajoo, Shyamal Lakhanpal, Abhishek Israni
arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.
By Koyar Afrasyab
arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.
By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
arXiv:2603. 24481v2 Announce Type: replace Abstract: Miscalibrated confidence scores are a practical obstacle to deploying AI in clinical settings.
By John Ray B. Martinez
arXiv:2607. 28608v1 Announce Type: new Abstract: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups.
By Sparsh Roy, Samuel Girmachew, Nishita Chavan
The paper introduces Evidence-Carrying Termination (ECT), a method that allows tool‑using large language model agents to declare a task complete only when a typed certificate links every required answer claim to valid, in‑scope trace evidence and a deterministic replay confirms the claimed value. In controlled experiments across 48 synthetic tasks and 576 trajectories, ECT eliminated unsafe completions and premature unsupported terminations, outperforming existing termination critics and controllers while maintaining comparable supported completion rates.
By Jason Liu
The paper introduces a new evaluation method called "same-input rerun" to assess the consistency of clinical language‑model agents across repeated runs. By replaying 1,000 MedAgentBench tasks with identical inputs, the authors find that action‑level outputs—such as test orders, medication requests, and referrals—vary significantly, even when benchmark scores remain unchanged. The study demonstrates that current benchmarks, which typically evaluate only a single run per task, can miss substantial behavioral divergence.
By Rohith Reddy Bellibatlu, Manpreet Singh, Zhoutian Han, Wenbin Zhang
The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.
By Melika Baghi
arXiv:2607. 15166v1 Announce Type: new Abstract: Most medical AI benchmarks measure whether a model knows the correct answer.
By Goktug Ozkan
arXiv:2607. 24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic coding, CPT for procedure billing, and HL7 FHIR for data exchange.
By Jianru Shen