arXiv AI By Sofiane Ouaari, Kevin Vorwalder, Nico Pfeifer

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

Read the original on arXiv AI →

arXiv:2607. 01973v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answering.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 7

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

The paper introduces the Medical Data Standardization Benchmark (MDS‑Bench), which evaluates vision‑language models (VLMs) on their ability to process raw, heterogeneous medical data. Models must identify source formats, convert raw images into VLM‑compatible inputs, extract relevant text, and organize the results into structured image‑text pairs. Experiments show that even the top VLM, Gemini 3 Flash, achieves only a 48.6% end‑to‑end success rate, underscoring the challenge of raw data standardization in clinical settings.

By Xin Chen, Dongliang Xu, Cunhao Zhu, Xudong Luo, Haoyang Lyu, Xiaoxiao Sun, Serena Yeung-Levy, Yue Yao
Hugging Face Trending Papers
Jul 6

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over given medical images and texts, implicitly assuming that standardized medical images, texts or question-answer pairs are already prepared. However, this assumption does not hold when we apply VLMs in real clinical practice, where medical data is often raw, heterogeneous, and fragmented across different sources.

arXiv Computer Vision
Aug 24

Toward Vision Language Model-based Assessment of Clinical Quality and Usability of LGE-MR Images for Cardiac Ablation Planning

This study introduces a two‑stage vision‑language model framework to assess the clinical quality and usability of late gadolinium enhancement (LGE) cardiac MRI images used for atrial fibrillation ablation planning. The first stage employs a fine‑tuned VLM to generate structured radiology‑style reports on five quality criteria—Noise, Motion Artifact, LA Boundary Accuracy, PV Region Accuracy, and Under‑segmentation Severity—while the second stage uses a GPT‑based reasoning module to convert these reports into structured quality scores and a binary decision on clinical usability. Evaluated on a curated dataset of 60 image‑slice and text‑pair annotations from 20 patients, the InternVL2 model achieved the highest criterion‑level accuracy, and DeepSeek reached perfect agreement on the clinical usability decision.

By Bipasha Kundu, Abhishek Chaturvedi, Axel W. E. Wismueller, Richard Simon, Cristian A. Linte