The paper introduces BRIE, a scalable framework that automatically creates question–answer pairs from longitudinal electronic health record notes, validated by nineteen clinicians. It offers a continuously maintainable benchmark for evaluating large language models in clinical settings, addressing limitations of manual, costly, and quickly outdated existing benchmarks. Experiments across nine LLMs and five inference strategies reveal that even state‑of‑the‑art systems often miss clinically important information, especially for synthesis‑heavy queries.
The paper introduces BRIE, a continuously maintainable benchmark for evaluating large language models (LLMs) in electronic health record (EHR) information retrieval. It presents a scalable framework that automatically generates question–answer pairs from longitudinal EHR notes, validated by nineteen clinicians. The benchmark allows assessment of multiple inference strategies and highlights that state‑of‑the‑art LLMs often miss clinically important information, especially when synthesis across documents is required.
By Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer
arXiv:2607. 06641v1 Announce Type: cross Abstract: Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hallucinations and the rapid evolution of official guidance.
By Felix Feldman, Joshua Harris, Timothy Laurence, Leo Loman, Ollie Higgins, Fan Grayson, Poonam Soma, Bethany Pace-Bonello, Michael Borowitz, Toby Nonnenmacher
arXiv:2606. 19266v1 Announce Type: cross Abstract: The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized domains and languages, yet the effectiveness of domain adaptation strategies remains unclear.
By Ikram Belmadani, Oumaima El Khettari, Carlos Ramisch, Frederic Bechet, Richard Dufour, Benoit Favre
MTDiag is a newly released multi-turn diagnostic dialogue dataset designed to evaluate large language models (LLMs) in clinically meaningful ways. It is built from DDXPlus, MIMIC-IV, and AJCR case reports, covering both common emergency department presentations and rare conditions, and normalizes cases into a canonical schema using UMLS concept identifiers and ICD-10 codes. The dataset includes a UserLM‑8B utterance‑generation pipeline and physician‑validated natural‑language utterances, and introduces clinical knowledge‑grounded metrics that go beyond simple diagnostic accuracy for multi‑turn differential diagnosis tasks.
By Pia Chouayfati, Alexander M. Fichtl, Miriam Ansch\"utz, George Doumat, Georg Groh
arXiv:2607. 24838v1 Announce Type: cross Abstract: In medical multiple-choice question answering (MCQA), Retrieval-Augmented Generation (RAG) can supplement the domain knowledge of language models (LMs).
By Seongwon Seo, Seung Hwan Cho, Young-Min Kim