arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.
By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv:2607. 02175v1 Announce Type: new Abstract: Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its "Hard" subset top score remains 32%.
By Samiha A. Ismail, Fan X. Chen, Ali Merali
arXiv:2608. 12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings.
By Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
arXiv:2609.12822v2 Announce Type: replace
Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs)....
By Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Jason A. Freed, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai
arXiv:2609.34024v1 Announce Type: cross
Abstract: Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and c...
By Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho
The study evaluates Jev 1.13, a non‑generative model that selects from predefined answer options, on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena‑MCQ, and the NEJM Case Challenges. Jev’s top‑1 accuracy matches GPT‑6 Sol with medium reasoning on PubMedQA but falls behind on MetaMedQA, DiagnosisArena‑MCQ, and NEJM cases. While Jev shows strong calibration on MetaMedQA and is fast and inexpensive, its performance on examination and complex diagnostic tasks is substantially lower, indicating the need for task‑specific validation before clinical deployment.
By Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho