arXiv Machine Learning By Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng, M. Tahar Kechadi

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

Read the original on arXiv Machine Learning →

arXiv:2608. 07550v1 Announce Type: cross Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

The study examines how differences in radiologists’ reporting styles—such as terminology, shorthand, formatting, and detail—affect the evaluation of AI-generated chest X‑ray reports. By quantifying the sensitivity of common metrics to these variations, the authors show that changes in reference reports can shift model rankings. They introduce a taxonomy of reporting variations and a rewriting method, ReRef, that preserves clinical meaning while altering style, and release a validated dataset of paired reference reports to aid future research.

By Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst
arXiv Machine Learning
Sep 21

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

The study evaluates the robustness of medical vision‑language models for tuberculosis screening on chest X‑rays by testing them across multiple datasets, prompts, and evaluation settings. Three specialized models (BioMedCLIP, CheXficient, MedSigLIP) and a general OpenCLIP model were audited on 12,200 images, producing 244,000 model–image–prompt scores. Results show that no model consistently outperforms others across all cohorts and reliability criteria, with prompt changes and control group composition significantly affecting AUROC, and that high training‑set performance does not reliably transfer to external cohorts.

By Mushir Akhtar, M. Tanveer, Mohd. Arshad