arXiv Machine Learning

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

arXiv:2608. 07550v1 Announce Type: cross Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment.

arXiv AI
Sep 17

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

The study examines how differences in radiologists’ reporting styles—such as terminology, shorthand, formatting, and detail—affect the evaluation of AI-generated chest X‑ray reports. By quantifying the sensitivity of common metrics to these variations, the authors show that changes in reference reports can shift model rankings. They introduce a taxonomy of reporting variations and a rewriting method, ReRef, that preserves clinical meaning while altering style, and release a validated dataset of paired reference reports to aid future research.

By Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst
arXiv Machine Learning
Sep 21

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

The study evaluates the robustness of medical vision‑language models for tuberculosis screening on chest X‑rays by testing them across multiple datasets, prompts, and evaluation settings. Three specialized models (BioMedCLIP, CheXficient, MedSigLIP) and a general OpenCLIP model were audited on 12,200 images, producing 244,000 model–image–prompt scores. Results show that no model consistently outperforms others across all cohorts and reliability criteria, with prompt changes and control group composition significantly affecting AUROC, and that high training‑set performance does not reliably transfer to external cohorts.

By Mushir Akhtar, M. Tanveer, Mohd. Arshad
arXiv AI
Aug 17

Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

arXiv:2608. 14399v1 Announce Type: cross Abstract: Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible.

By Syeda Anshrah Gillani, Mirza Samad Ahmed Baig
arXiv AI
2d ago

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.

By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang