arXiv AI

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

arXiv:2607. 01973v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answering.

arXiv Computer Vision
Sep 7

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

The paper introduces the Medical Data Standardization Benchmark (MDS‑Bench), which evaluates vision‑language models (VLMs) on their ability to process raw, heterogeneous medical data. Models must identify source formats, convert raw images into VLM‑compatible inputs, extract relevant text, and organize the results into structured image‑text pairs. Experiments show that even the top VLM, Gemini 3 Flash, achieves only a 48.6% end‑to‑end success rate, underscoring the challenge of raw data standardization in clinical settings.

By Xin Chen, Dongliang Xu, Cunhao Zhu, Xudong Luo, Haoyang Lyu, Xiaoxiao Sun, Serena Yeung-Levy, Yue Yao
Hugging Face Trending Papers
Jul 6

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over given medical images and texts, implicitly assuming that standardized medical images, texts or question-answer pairs are already prepared. However, this assumption does not hold when we apply VLMs in real clinical practice, where medical data is often raw, heterogeneous, and fragmented across different sources.

arXiv Computer Vision
Aug 24

Toward Vision Language Model-based Assessment of Clinical Quality and Usability of LGE-MR Images for Cardiac Ablation Planning

This study introduces a two‑stage vision‑language model framework to assess the clinical quality and usability of late gadolinium enhancement (LGE) cardiac MRI images used for atrial fibrillation ablation planning. The first stage employs a fine‑tuned VLM to generate structured radiology‑style reports on five quality criteria—Noise, Motion Artifact, LA Boundary Accuracy, PV Region Accuracy, and Under‑segmentation Severity—while the second stage uses a GPT‑based reasoning module to convert these reports into structured quality scores and a binary decision on clinical usability. Evaluated on a curated dataset of 60 image‑slice and text‑pair annotations from 20 patients, the InternVL2 model achieved the highest criterion‑level accuracy, and DeepSeek reached perfect agreement on the clinical usability decision.

By Bipasha Kundu, Abhishek Chaturvedi, Axel W. E. Wismueller, Richard Simon, Cristian A. Linte
Hugging Face Trending Papers
Jun 29

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness. However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.

arXiv Machine Learning
Jul 30

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.

By Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Sch\"onlieb, Anna Korhonen, Anna Breger
arXiv AI
2d ago

VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision

VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.

By Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen
arXiv Machine Learning
Jun 26

Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA

arXiv:2606. 27023v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) applied to Medical Visual Question Answering (VQA) tend to produce overconfident outputs regardless of actual correctness, and existing verbalized confidence calibration methods, developed primarily for text only LLMs, do not account for the multimodal nature of medical image understanding.

By Eren Senoglu, Federico Toschi, Nicolo Brunello, Andrea Sassella, Mark James Carman