arXiv:2605. 03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer.
By Jingxi Qiu, Zeyu Han, Cheng Huang
arXiv:2607. 20950v1 Announce Type: new Abstract: BoN improves model outputs by sampling several candidates and selecting one with a proxy score, but it assumes that complete candidates can be evaluated reliably.
By Cenwei Zhang, Teng Fang, Yuxia Wang, Derek Li, Bryan Dai, Lei You
arXiv:2609.12884v1 Announce Type: new
Abstract: A medical claim's correctness often depends not on the claim alone, but on the clinical structure around it. A claim may require a lab reference range,...
By Hasan Iqbal, Sarfraz Ahmad, Hyunjae Kim, Sihyeon Park, Junjie Liao, Qingyu Chen, Preslav Nakov, Yuxia Wang
arXiv:2605. 28910v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect statements that limit their reliability in specialized healthcare applications.
By Shamanth Kuthpadi Seethakantha, Dung Ngoc Thai, Vara Prasad Gudi, Simran Tiwari, Rami Matar, Avijit Mitra, Wenlong Zhao, Andrew McCallum, Wael Salloum
arXiv:2609.18852v1 Announce Type: new
Abstract: Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed t...
By Fengnan Li, Heman Burre, Liwen Sun, Roshni Varma, Matthew M. Engelhard
BoN improves model outputs by sampling several candidates and selecting one with a proxy score, but it assumes that complete candidates can be evaluated reliably. Many vision-language tasks instead provide only partial verification: a finding, span, value, region, or relation may be checkable even when no dependable whole-response verifier exists.
The paper investigates why retrieval‑based open‑ended evaluation fails in medical fact verification. By creating two detailed taxonomies—one for retrieval‑stage errors across five quality dimensions and another for verifier‑reasoning errors across six steps—the authors automatically label evidence quality and reasoning errors using an LLM‑as‑Judge pipeline. Their large‑scale stress tests across multiple retrieval methods and verifier models show that increasing model size, reasoning effort, source breadth, or medical fine‑tuning does not eliminate these failure modes, indicating fundamental limits of the retrieve‑then‑verify paradigm in open‑ended medical contexts.
By Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi, Jie Gao, Bernal Jim\'enez Guti\'errez, Mark Dredze
arXiv:2607. 18999v1 Announce Type: cross Abstract: Multi-turn medical consultation agents must decide what to ask, adapt to patient responses, and determine when the collected evidence is sufficient.
By Guofeng Zhang, Yizeng Quan, Huaiyi Fang, Jianwei Lv, Jinyao Liu, Xunxu Duan, Lening An, Yu Ouyang, Junfeng Wang
arXiv:2608.31016v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2606. 05970v1 Announce Type: cross Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks.
By Martin Murin
arXiv:2607. 28788v1 Announce Type: new Abstract: Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence.
By Jiahui Li, Ruili Fang, Zishuai Liu, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou
arXiv:2609.12822v2 Announce Type: replace
Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs)....
By Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Jason A. Freed, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai