arXiv AI

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Med-AR introduces two autoregressive vision‑language models, Med‑AR‑8B and Med‑AR‑2B, pretrained on structured radiology reports, abnormality‑focused text, and region annotations to address long‑tailed chest X‑ray classification. The models outperform existing contrastive, self‑supervised, and supervised encoders—including Med‑CLIP, CheXFound, EVA‑Base, ARK, and BioViL‑T—across PadChest, MIMIC‑CXR, and CheXpert, achieving higher mean AUROC and AUPRC for head, medium, and tail findings and lower excess area under the risk‑coverage curve. Med‑AR also demonstrates improved selective‑prediction performance, with Med‑AR‑8B raising tail‑label mean AUPRC on MIMIC‑CXR from 0.1033 to 0.1441 and Med‑AR‑2B delivering the strongest discrimination on PadChest.

arXiv AI
Aug 5

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.

By Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
arXiv Computer Vision
Aug 27

Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

The paper introduces a label‑free method called AURCC for selecting the best foundational model for medical image classification when the target domain lacks labels. AURCC uses a pseudo‑label discrepancy computed by the SUDO framework to score models without fine‑tuning. Experiments on chest X‑ray data across three inter‑hospital shifts show that AURCC closely matches the true model ranking, outperforming simple source‑accuracy baselines especially when source data are limited.

By Juan I\~naki Larrea, Lucas Mansilla, Enzo Ferrante
arXiv Machine Learning
Jul 30

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.

By Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Sch\"onlieb, Anna Korhonen, Anna Breger
arXiv Computer Vision
Sep 22

CXR-LT 2026 Challenge: Multi-Center Long-Tailed and Zero Shot Chest X-ray Classification

The CXR‑LT 2026 Challenge introduces a multi‑center, long‑tailed chest X‑ray classification benchmark with over 145,000 radiologist‑annotated images from PadChest and NIH datasets. It defines two core tasks: robust multi‑label classification on 30 known classes and open‑world generalization to 6 unseen rare disease classes. The paper outlines data collection, annotation, solution strategies, and evaluates performance across head‑vs‑tail, calibration, and cross‑center gaps, noting that vision‑language models improve in‑distribution and zero‑shot performance but rare‑finding detection under multi‑center shift remains difficult.

By Hexin Dong, Yi Lin, Pengyu Zhou, Fengnian Zhao, Alan Clint Legasto, Juno Cho, Dohui Kim, Justin Namuk Kim, Mingeon Kim, Sunwoo Kwak, Gabriel Moy\`a-Alcover, Ky Trung Nguyen, Thanh-Huy Nguyen, Ha-Hieu Pham, Huy-Hieu Pham, Huy Le Pham, Nikhileswara Rao Sulake, Aina Tur-Serrano, Ruichi Zhang, Ang Zu, Adam E. Flanders, Zhiyong Lu, Ronald M. Summers, Mingquan Lin, Hao Chen, Yuzhe Yang, George Shih, Yifan Peng
arXiv Machine Learning
Aug 4

RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding

arXiv:2608. 00147v1 Announce Type: cross Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility.

By Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu, Miriam Kumpf, Lena Schmitzer, Lea Schumann, Jannik Kahmann, Friedrich Puttkammer, Johannes Moll, Jannik L\"ubberstedt, Zeineb Ben Chaaben, Anirudh Narayanan, Cosmin I. Bercea, Sebastian Ziegelmayer, Marcus R. Makowski, Daniel Rueckert, Lisa C. Adams, Keno K. Bressem
Hugging Face Trending Papers
Jun 29

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness. However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.

arXiv Machine Learning
Sep 21

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

The study evaluates the robustness of medical vision‑language models for tuberculosis screening on chest X‑rays by testing them across multiple datasets, prompts, and evaluation settings. Three specialized models (BioMedCLIP, CheXficient, MedSigLIP) and a general OpenCLIP model were audited on 12,200 images, producing 244,000 model–image–prompt scores. Results show that no model consistently outperforms others across all cohorts and reliability criteria, with prompt changes and control group composition significantly affecting AUROC, and that high training‑set performance does not reliably transfer to external cohorts.

By Mushir Akhtar, M. Tanveer, Mohd. Arshad
arXiv Computer Vision
Aug 27

What Do Medical Vision-Language Models Learn in Radiology? Transfer, Alignment, and Source-Proxy Leakage Under Distribution Shift

The paper investigates how medical vision‑language models (VLMs) behave when faced with distribution shifts such as changes in acquisition domain, supervision, or evaluation protocol. Using datasets like NIH ChestXray14, CheXpert, PadChest, and OpenI, the authors isolate cross‑dataset visual transfer, evaluate multimodal alignment, and quantify source‑proxy leakage in frozen embeddings. They find that self‑supervised visual initialization improves transfer, adversarial adaptation is only marginally helpful, and that multimodal retrieval performance drops under external stress tests while source‑proxy information remains recoverable, highlighting hidden failure modes in medical VLMs.

By Ayoub Louaye Bouaziz, Lokmane Chebouba, Yassine Himeur