arXiv Computation and Language

Evaluating VQA in Vision Language Models using Cooperative Principles

The paper evaluates Vision Language Models (VLMs) on Visual Question Answering tasks where questions violate Grice's maxims. By generating question modifiers that add non-essential, ambiguous, or false information, the authors show that VLMs such as ChatGPT, Claude, Gemini, and Llava exhibit reduced performance. They also compare human pragmatic reasoning to VLM reasoning, noting differences in how each handles human‑induced versus AI‑generated violations, and find that humans spend less time resolving VLM‑induced violations while VLMs are less accurate in those cases.

arXiv Computer Vision
Sep 23

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.

By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao
arXiv AI
5d ago

Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning

Hob‑VL is a benchmark for visually grounded Boolean reasoning that tests models on two tasks: determining whether a Boolean rule holds in an image and identifying the unique object that satisfies a Boolean description. It contains 6,000 balanced Yes/No questions built from ten visual statements across 1,000 generated scenes and 46 labeled photographs, plus 1,000 object‑identification questions on the same photographs. The questions are designed to challenge reasoning with misleading local cues, nested logical operations, and both symbolic and natural‑language presentations, revealing significant gaps in current model performance.

By Yuzhou Wang, Emile Anand, Ijay Narang
arXiv AI
Aug 10

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.

By Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy
arXiv AI
Sep 18

Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

The paper shows that the wording of prompts in vision‑language models (VLMs) can either improve or worsen robustness to image corruption. Verbose prompts broaden the cross‑modal attention’s frequency filter, making the model less sensitive to corruptions, while semantically complex prompts narrow the filter and increase vulnerability. Experiments on Qwen3‑VL and LLaVA‑OneVision confirm that adding padding or verbose phrasing reduces answer drift by 70–81% on 8B models.

By Farooq Ahmad Wani, Maria Sofia Bucarelli, Mujtaba Hussain Mirza, Oleksandr Pryymak, Aryo Pradipta Gema, Iacopo Masi, Pasquale Minervini, Fabrizio Silvestri
arXiv Computation and Language
Sep 7

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

MedProb is a lightweight probing framework that predicts multiple-choice medical visual question answering (Med‑VQA) answers directly from frozen vision‑language model (VLM) representations, avoiding free‑text generation. On datasets such as PATH‑VQA, SLAKE, and VQA‑RAD, MedProb extracts more answer‑relevant signal than prompting and outperforms both medical VLMs and agentic systems. The approach also narrows the performance gap between small and large models, shows that medical adaptation does not consistently improve linear decodability, and reveals positional biases in both prompting and generation.

By Erfan Nourbakhsh, Ke Yang, Anthony Rios
arXiv Machine Learning
Jun 15

Self-Evolving Visual Questioner

arXiv:2606. 13929v1 Announce Type: cross Abstract: Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains underexplored.

By Yijun Liang, Hengguang Zhou, Ming Li, Lichen Li, Cho-Jui Hsieh, Tianyi Zhou
arXiv AI
Sep 15

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

NoteVQA is a new benchmark that collects 252 real‑life visual questions from the Chinese image‑sharing platform Xiaohongshu, covering 12 topics and 7 user intents. Each question is paired with a concise expert reference and a human‑audited interleaved answer that blends text and visual evidence. The study evaluates VLMs on short‑answer correctness and interleaved answer quality using a new AgenticInterleave framework and a 12‑dimensional IVR‑12 rubric, finding that even state‑of‑the‑art models achieve only about 53% accuracy and lag behind human references in content quality.

By Haonan Jiang, Guojian Zhan, Jiancong Xie, Shijun Wan, Dongiia Zhao, Cheng Chen, Yahui Liu, Yao Hu, Chuan Mu