arXiv:2609.01470v1 Announce Type: new
Abstract: As AI systems are increasingly used to draft radiology reports, reliably evaluating their clinical quality remains a critical challenge. Large language...
By Charles Corbi\`ere, L\'eo Machado, Aubin Charley, Baptiste Callard, Pierre Manceron, Corentin Dancette
arXiv:2606. 17062v1 Announce Type: cross Abstract: Radiology report evaluation must distinguish clinical compatibility from surface similarity, because negation, laterality, or normal-abnormal polarity can reverse a finding.
By Zhenhong Yang, Zhuoyun Liu, Jintao Fei, Wen Tang, Shichao Quan, Jun Zhao, Jun Xu
arXiv:2606. 08769v1 Announce Type: cross Abstract: Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polarity reversals, location changes, uncertainty mismatches, and temporal-comparison errors rather than low surface similarity alone.
By Weixin Liu, Juming Xiong, Yang Li, Qingyuan Song, Susannah Rose, Murat Kantarcioglu, Bradley Malin, Zhijun Yin
The study examines how differences in radiologists’ reporting styles—such as terminology, shorthand, formatting, and detail—affect the evaluation of AI-generated chest X‑ray reports. By quantifying the sensitivity of common metrics to these variations, the authors show that changes in reference reports can shift model rankings. They introduce a taxonomy of reporting variations and a rewriting method, ReRef, that preserves clinical meaning while altering style, and release a validated dataset of paired reference reports to aid future research.
By Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst
arXiv:2608.31016v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.
By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
The study evaluates Jev 1.13, a non‑generative model that selects from predefined answer options, on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena‑MCQ, and the NEJM Case Challenges. Jev’s top‑1 accuracy matches GPT‑6 Sol with medium reasoning on PubMedQA but falls behind on MetaMedQA, DiagnosisArena‑MCQ, and NEJM cases. While Jev shows strong calibration on MetaMedQA and is fast and inexpensive, its performance on examination and complex diagnostic tasks is substantially lower, indicating the need for task‑specific validation before clinical deployment.
By Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho
arXiv:2607. 05880v1 Announce Type: cross Abstract: Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone.
By Suneeta Mall, Vladimir Nekrasov, Ashnil Kumar, Sajith Karunasena, Aiden Nibali, Alix Bird, Mateo Diaz Shine, Jarrel Seah
arXiv:2609.34024v1 Announce Type: cross
Abstract: Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and c...
By Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho
The paper investigates why retrieval‑based open‑ended evaluation fails in medical fact verification. By creating two detailed taxonomies—one for retrieval‑stage errors across five quality dimensions and another for verifier‑reasoning errors across six steps—the authors automatically label evidence quality and reasoning errors using an LLM‑as‑Judge pipeline. Their large‑scale stress tests across multiple retrieval methods and verifier models show that increasing model size, reasoning effort, source breadth, or medical fine‑tuning does not eliminate these failure modes, indicating fundamental limits of the retrieve‑then‑verify paradigm in open‑ended medical contexts.
By Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi, Jie Gao, Bernal Jim\'enez Guti\'errez, Mark Dredze
KnowBench is a new benchmark for clinical AI that measures Effort Reduction (ER), the proportion of system-generated clinical work product accepted by clinicians after expert and safety review. The metric is applied uniformly across various administrative tasks—visit notes, billing codes, orders, EHR summarization, patient summaries, and decision support—using the clinician’s review-and-attestation as ground truth. An initial deployment of Knowtex’s models achieved an aggregate ER of 97.99% across more than one million encounters in six months, with specialty-specific ER ranging from 96.8% to 98.9%.
By Jocelyn Kang, Caroline Zhang
arXiv:2609.12822v2 Announce Type: replace
Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs)....
By Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Jason A. Freed, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai