arXiv:2606. 00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions common in clinical practice.
By Andrei Marian Feier, Veysel Kocaman, Yigit Gul, Ahmet Korkmaz, Alexander Thomas, Aleksei Zakharov, Jay Gil, Mehmet Butgul, David Talby
arXiv:2508. 00923v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against.
By Jiazhen Pan (Cherise), Bailiang Jian (Cherise), Paul Hager (Cherise), Yundi Zhang (Cherise), Che Liu (Cherise), Friederike Jungmann (Cherise), Hongwei Bran Li (Cherise), Julian Canisius (Cherise), Chenyu You (Cherise), Junde Wu (Cherise), Jiayuan Zhu (Cherise), Fenglin Liu (Cherise), Yuyuan Liu (Cherise), Niklas Bubeck (Cherise), Moritz Knolle (Cherise), Chen (Cherise), Chen (Cherise), Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert
arXiv:2608. 16643v1 Announce Type: cross Abstract: Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation.
By Yifan Zhang, Rahmatollah Beheshti
arXiv:2604.26766v2 Announce Type: replace-cross
Abstract: Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly varia...
By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart.
arXiv:2607. 20462v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking.
By Melanie Rieff, Robin Staab, Thibaud Gloaguen, Stefan Hegselmann, Martin Vechev
The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.
By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
arXiv:2608.29582v1 Announce Type: cross
Abstract: Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigati...
By Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi
The paper investigates how medical large language models (LLMs) may exhibit narrative anchoring bias when presented with the same clinical case in different patient voices. Using the NarrativeShield SDoH MedQA dataset, the authors evaluate three Qwen2.5 instruction‑tuned LLMs (1.5B, 3B, 7B) on 300 clinical cases, reporting metrics such as persona‑level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. The 7B model achieves the highest accuracy (56.33 %) and correct consistency (40.33 %), yet narrative sensitivity errors remain substantial (31.67 %).
By Ahnaf Atef Choudhury, Ramkrishna Saha
arXiv:2607. 14385v1 Announce Type: cross Abstract: Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requiring different interventions.
By Thanni Adewuyi, Anuoluwa Sotome, Samuel Okoko, Angel Ezendu, Oluwafunke Akinbuwa, Oluwaseun Odunsi, Oluwasegun Oguntuase, Oluwadarasimi Oguntuase, Ifeoma Nwabueze, Abiodun Adereni
arXiv:2606. 07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summarization.
By Mahdi Alkaeed
arXiv:2602.06268v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly integrated into clinical workflows. However, p...
By Junhyeok Lee, Han Jang, Kyu Sung Choi