arXiv:2603.11413v4 Announce Type: replace-cross
Abstract: A recent Nature Medicine study reported that ChatGPT Health under-triages 51.6% of emergencies and concluded that consumer-facing AI triage p...
By David Fraile Navarro, Jialei Sheng, Farah Magrabi, Enrico Coiera
arXiv:2607. 18086v1 Announce Type: new Abstract: Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved.
By Koyar Afrasyab
arXiv:2608.31016v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2604. 07709v4 Announce Type: replace-cross Abstract: A heavily safety-trained model will hand a physician the full, patient-followable benzodiazepine taper and refuse it to the patient who needs it, over identical clinical facts; the knowledge is present either way.
By David Gringras
IatroBench is a pre‑registered benchmark that evaluates language models on clinical omission and commission harms across 60 scenarios and six models. Using a physician‑written rubric scored by Claude Opus 4.6, the study finds that models tend to withhold more information from patients than from doctors—a phenomenon termed framing‑contingent withholding—while also revealing varied patterns of omission across different models. The benchmark highlights how framing influences the amount of medical information shared by AI systems.
By David Gringras
arXiv:2608.31017v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 1...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2608.23397v1 Announce Type: new
Abstract: Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final diagnosis a...
By Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao
arXiv:2606. 03198v1 Announce Type: cross Abstract: Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluation conditions has not been quantitatively characterized.
By Sangwon Baek, Kyu Yeon Hur, Kyunga Kim
arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.
By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
arXiv:2607. 09804v1 Announce Type: cross Abstract: Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications.
By Avi-ad Avraam Buskila
The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.
By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
arXiv:2608. 02617v1 Announce Type: cross Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings.
By Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley