arXiv AI

Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts

arXiv:2608. 15254v1 Announce Type: new Abstract: Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI).

arXiv AI
2d ago

Scaling Clinical Judgment to Evaluate Medical AI

arXiv:2609.12822v2 Announce Type: replace Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs)....

By Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Jason A. Freed, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai
arXiv AI
Aug 17

Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

arXiv:2608. 14399v1 Announce Type: cross Abstract: Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible.

By Syeda Anshrah Gillani, Mirza Samad Ahmed Baig
arXiv AI
Sep 25

IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models

IatroBench is a pre‑registered benchmark that evaluates language models on clinical omission and commission harms across 60 scenarios and six models. Using a physician‑written rubric scored by Claude Opus 4.6, the study finds that models tend to withhold more information from patients than from doctors—a phenomenon termed framing‑contingent withholding—while also revealing varied patterns of omission across different models. The benchmark highlights how framing influences the amount of medical information shared by AI systems.

By David Gringras
arXiv Computation and Language
Aug 24

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

The study evaluates large language models (LLMs) on unprocessed electronic medical record data for clinical registry abstraction, focusing on the American College of Cardiology National Cardiovascular Data Registry. In a pilot at one academic center, the LLM identified candidate data sources for each registry question, which abstractors used to define question‑specific document sets. In a subsequent validation at a second center, the LLM answered 157 registry questions with an overall mean accuracy of 91.5%, but accuracy dropped from 96% for simple medication or event flag questions to 62% for event timing questions, reflecting increasing ambiguity and required clinical reasoning.

By James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker