arXiv AI By Koyar Afrasyab

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

Read the original on arXiv AI →

arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models

IatroBench is a pre‑registered benchmark that evaluates language models on clinical omission and commission harms across 60 scenarios and six models. Using a physician‑written rubric scored by Claude Opus 4.6, the study finds that models tend to withhold more information from patients than from doctors—a phenomenon termed framing‑contingent withholding—while also revealing varied patterns of omission across different models. The benchmark highlights how framing influences the amount of medical information shared by AI systems.

By David Gringras