arXiv:2607. 18086v1 Announce Type: new Abstract: Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved.
By Koyar Afrasyab
arXiv:2608.31016v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2604. 07709v4 Announce Type: replace-cross Abstract: A heavily safety-trained model will hand a physician the full, patient-followable benzodiazepine taper and refuse it to the patient who needs it, over identical clinical facts; the knowledge is present either way.
By David Gringras
IatroBench is a pre‑registered benchmark that evaluates language models on clinical omission and commission harms across 60 scenarios and six models. Using a physician‑written rubric scored by Claude Opus 4.6, the study finds that models tend to withhold more information from patients than from doctors—a phenomenon termed framing‑contingent withholding—while also revealing varied patterns of omission across different models. The benchmark highlights how framing influences the amount of medical information shared by AI systems.
By David Gringras
arXiv:2608.31017v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 1...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2608.23397v1 Announce Type: new
Abstract: Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final diagnosis a...
By Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao