arXiv Computer Vision
Sep 23

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.

By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao
arXiv AI
Aug 14

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).

By Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
arXiv Computer Vision
Sep 22

Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning

The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.

By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha