In clinical practice, agreement among independent experts is treated as evidence of reliability, and multi-round consensus has become a core mechanism of agentic medical question-answering systems. Wh...
arXiv:2603. 24481v2 Announce Type: replace Abstract: Miscalibrated confidence scores are a practical obstacle to deploying AI in clinical settings.
By John Ray B. Martinez
arXiv:2605. 03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer.
By Jingxi Qiu, Zeyu Han, Cheng Huang
arXiv:2609.08016v1 Announce Type: new
Abstract: Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagre...
By Chen Qian
arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.
By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
arXiv:2608. 07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably.
By Veronica Chatrath, Bryan Zhu, George Pu, Jingxuan Fan, Apaar Shanker, Varun Ursekar, Anahita Sharma, Jason Qin, Keqi Han, Soham Dinesh Tiwari, Soham Dan, Vijay Kalmath, Yuan Li, Daniel Yue Zhang, Chenguang Wang, Zainab Doctor, Zhijun Yin, Nigam H. Shah, Yuan Xue
The study evaluates whether the chain-of-thought (CoT) rationales produced by medical language models truly influence their answers. Using a 30‑operator perturbation audit that modifies both the question and the CoT (e.g., severity reversal, negation flip, demographic swap, evidence ablation), the authors found that 72.9% of edits did not change the model’s answer—a high Chain‑Decoupling Rate (CDR). Across 14 large language models and four medical QA benchmarks, the CoT text had little impact on accuracy, and removing CoT prompting did not reduce performance.
"whyItMatters":"The findings suggest that current medical CoT outputs may be more documentation than genuine reasoning, highlighting the need for better faithfulness checks in clinical AI systems."
By Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long
The paper introduces a pluralistic agreement index, Gamma, to quantify how often wrong runs of large language models (LLMs) agree with the majority consensus. By decomposing Gamma into a mechanical component and a preference‑unexplained residual, the authors show that on GPT‑4.1 the mechanical part explains most of the agreement on multiple‑choice benchmarks but only about half on open‑domain tasks, revealing a residual bias that can cause self‑consistency to backfire on hard questions. The study provides a quantitative framework for understanding when majority voting over LLM samples improves or harms accuracy, without proposing new voting methods.
By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.
By Koyar Afrasyab
arXiv:2609.34024v1 Announce Type: cross
Abstract: Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and c...
By Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho
The study evaluates Jev 1.13, a non‑generative model that selects from predefined answer options, on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena‑MCQ, and the NEJM Case Challenges. Jev’s top‑1 accuracy matches GPT‑6 Sol with medium reasoning on PubMedQA but falls behind on MetaMedQA, DiagnosisArena‑MCQ, and NEJM cases. While Jev shows strong calibration on MetaMedQA and is fast and inexpensive, its performance on examination and complex diagnostic tasks is substantially lower, indicating the need for task‑specific validation before clinical deployment.
By Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho
The study evaluates large language models (LLMs) on unprocessed electronic medical record data for clinical registry abstraction, focusing on the American College of Cardiology National Cardiovascular Data Registry. In a pilot at one academic center, the LLM identified candidate data sources for each registry question, which abstractors used to define question‑specific document sets. In a subsequent validation at a second center, the LLM answered 157 registry questions with an overall mean accuracy of 91.5%, but accuracy dropped from 96% for simple medication or event flag questions to 62% for event timing questions, reflecting increasing ambiguity and required clinical reasoning.
By James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker