arXiv AI

Cross-modal linkage risk in clinical vision-language models

arXiv:2606. 02276v1 Announce Type: cross Abstract: Vision-language models (VLMs) trained on paired chest radiographs and radiology reports learn a shared embedding space that can preserve instance-level image-report correspondence.

arXiv Computer Vision
Aug 27

What Do Medical Vision-Language Models Learn in Radiology? Transfer, Alignment, and Source-Proxy Leakage Under Distribution Shift

The paper investigates how medical vision‑language models (VLMs) behave when faced with distribution shifts such as changes in acquisition domain, supervision, or evaluation protocol. Using datasets like NIH ChestXray14, CheXpert, PadChest, and OpenI, the authors isolate cross‑dataset visual transfer, evaluate multimodal alignment, and quantify source‑proxy leakage in frozen embeddings. They find that self‑supervised visual initialization improves transfer, adversarial adaptation is only marginally helpful, and that multimodal retrieval performance drops under external stress tests while source‑proxy information remains recoverable, highlighting hidden failure modes in medical VLMs.

By Ayoub Louaye Bouaziz, Lokmane Chebouba, Yassine Himeur
arXiv AI
Sep 25

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Med-AR introduces two autoregressive vision‑language models, Med‑AR‑8B and Med‑AR‑2B, pretrained on structured radiology reports, abnormality‑focused text, and region annotations to address long‑tailed chest X‑ray classification. The models outperform existing contrastive, self‑supervised, and supervised encoders—including Med‑CLIP, CheXFound, EVA‑Base, ARK, and BioViL‑T—across PadChest, MIMIC‑CXR, and CheXpert, achieving higher mean AUROC and AUPRC for head, medium, and tail findings and lower excess area under the risk‑coverage curve. Med‑AR also demonstrates improved selective‑prediction performance, with Med‑AR‑8B raising tail‑label mean AUPRC on MIMIC‑CXR from 0.1033 to 0.1441 and Med‑AR‑2B delivering the strongest discrimination on PadChest.

By Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, Preetham Putha
arXiv Machine Learning
Aug 4

RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding

arXiv:2608. 00147v1 Announce Type: cross Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility.

By Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu, Miriam Kumpf, Lena Schmitzer, Lea Schumann, Jannik Kahmann, Friedrich Puttkammer, Johannes Moll, Jannik L\"ubberstedt, Zeineb Ben Chaaben, Anirudh Narayanan, Cosmin I. Bercea, Sebastian Ziegelmayer, Marcus R. Makowski, Daniel Rueckert, Lisa C. Adams, Keno K. Bressem
Hugging Face Trending Papers
Jun 29

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness. However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.

arXiv AI
Jul 23

Self-supervision drives representational convergence in medical foundation models more than clinical supervision

arXiv:2607. 20274v1 Announce Type: cross Abstract: Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure.

By Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung, Jakob Nikolas Kather, Daniel Truhn
arXiv AI
Sep 15

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

arXiv:2609.15635v1 Announce Type: cross Abstract: A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a pai...

By Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
arXiv AI
Sep 3

Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts

The paper evaluates federated learning with Low‑Rank Adaptation (LoRA) for fine‑tuning the BiomedCLIP vision‑language model on chest X‑ray classification across four international cohorts. Federated LoRA improves shared‑class AUC from 0.687 to 0.802, outperforming isolated single‑cohort training and approaching a centralized reference. The study shows that SVD‑based product‑space aggregation (FlexLoRA) is crucial for performance, while FedProx offers no advantage over FedAvg in this setting.

By Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal, Debesh Jha, Sunil Kumar Gaire
arXiv Machine Learning
Sep 21

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

The study evaluates the robustness of medical vision‑language models for tuberculosis screening on chest X‑rays by testing them across multiple datasets, prompts, and evaluation settings. Three specialized models (BioMedCLIP, CheXficient, MedSigLIP) and a general OpenCLIP model were audited on 12,200 images, producing 244,000 model–image–prompt scores. Results show that no model consistently outperforms others across all cohorts and reliability criteria, with prompt changes and control group composition significantly affecting AUROC, and that high training‑set performance does not reliably transfer to external cohorts.

By Mushir Akhtar, M. Tanveer, Mohd. Arshad