arXiv AI

It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them

arXiv AI
Aug 14

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).

By Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
arXiv AI
5d ago

The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties

The paper demonstrates that vision‑language models’ refusal behavior is heavily influenced by whether an image is attached to a request, even when the image is blank or unreadable. Attaching such an image shifts refusal scores by large margins for borderline‑benign prompts while leaving genuinely neutral instructions largely unchanged. This effect varies with image properties, persists across checkpoints, and is not mitigated by explicit instructions to ignore the image.

By Haoyu Zhang, Yi Feng, Hanwen Liu, Shibo Zheng, Zhuoxi Wang, Yang Chen, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv AI
Sep 18

Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

The paper shows that the wording of prompts in vision‑language models (VLMs) can either improve or worsen robustness to image corruption. Verbose prompts broaden the cross‑modal attention’s frequency filter, making the model less sensitive to corruptions, while semantically complex prompts narrow the filter and increase vulnerability. Experiments on Qwen3‑VL and LLaVA‑OneVision confirm that adding padding or verbose phrasing reduces answer drift by 70–81% on 8B models.

By Farooq Ahmad Wani, Maria Sofia Bucarelli, Mujtaba Hussain Mirza, Oleksandr Pryymak, Aryo Pradipta Gema, Iacopo Masi, Pasquale Minervini, Fabrizio Silvestri
arXiv AI
Jul 29

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li
arXiv AI
1d ago

VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision

VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.

By Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen
arXiv AI
Sep 15

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

arXiv:2609.15635v1 Announce Type: cross Abstract: A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a pai...

By Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
arXiv Computer Vision
Sep 23

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.

By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao