arXiv AI By Mateusz Koz{\l}owski

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

Read the original on arXiv AI →

arXiv:2607. 25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 21

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

The study evaluates the robustness of medical vision‑language models for tuberculosis screening on chest X‑rays by testing them across multiple datasets, prompts, and evaluation settings. Three specialized models (BioMedCLIP, CheXficient, MedSigLIP) and a general OpenCLIP model were audited on 12,200 images, producing 244,000 model–image–prompt scores. Results show that no model consistently outperforms others across all cohorts and reliability criteria, with prompt changes and control group composition significantly affecting AUROC, and that high training‑set performance does not reliably transfer to external cohorts.

By Mushir Akhtar, M. Tanveer, Mohd. Arshad
arXiv AI
Sep 17

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

The study examines how differences in radiologists’ reporting styles—such as terminology, shorthand, formatting, and detail—affect the evaluation of AI-generated chest X‑ray reports. By quantifying the sensitivity of common metrics to these variations, the authors show that changes in reference reports can shift model rankings. They introduce a taxonomy of reporting variations and a rewriting method, ReRef, that preserves clinical meaning while altering style, and release a validated dataset of paired reference reports to aid future research.

By Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst
arXiv AI
Jul 8

Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context

arXiv:2607. 05880v1 Announce Type: cross Abstract: Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone.

By Suneeta Mall, Vladimir Nekrasov, Ashnil Kumar, Sajith Karunasena, Aiden Nibali, Alix Bird, Mateo Diaz Shine, Jarrel Seah
arXiv Machine Learning
Aug 11

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

arXiv:2608. 07550v1 Announce Type: cross Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment.

By Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng, M. Tahar Kechadi
arXiv AI
Sep 25

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Med-AR introduces two autoregressive vision‑language models, Med‑AR‑8B and Med‑AR‑2B, pretrained on structured radiology reports, abnormality‑focused text, and region annotations to address long‑tailed chest X‑ray classification. The models outperform existing contrastive, self‑supervised, and supervised encoders—including Med‑CLIP, CheXFound, EVA‑Base, ARK, and BioViL‑T—across PadChest, MIMIC‑CXR, and CheXpert, achieving higher mean AUROC and AUPRC for head, medium, and tail findings and lower excess area under the risk‑coverage curve. Med‑AR also demonstrates improved selective‑prediction performance, with Med‑AR‑8B raising tail‑label mean AUPRC on MIMIC‑CXR from 0.1033 to 0.1441 and Med‑AR‑2B delivering the strongest discrimination on PadChest.

By Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, Preetham Putha