arXiv Computation and Language By Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal

Evaluating VQA in Vision Language Models using Cooperative Principles

Read the original on arXiv Computation and Language →

The paper evaluates Vision Language Models (VLMs) on Visual Question Answering tasks where questions violate Grice's maxims. By generating question modifiers that add non-essential, ambiguous, or false information, the authors show that VLMs such as ChatGPT, Claude, Gemini, and Llava exhibit reduced performance. They also compare human pragmatic reasoning to VLM reasoning, noting differences in how each handles human‑induced versus AI‑generated violations, and find that humans spend less time resolving VLM‑induced violations while VLMs are less accurate in those cases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computer Vision
Sep 23

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.

By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao
arXiv AI
5d ago

Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning

Hob‑VL is a benchmark for visually grounded Boolean reasoning that tests models on two tasks: determining whether a Boolean rule holds in an image and identifying the unique object that satisfies a Boolean description. It contains 6,000 balanced Yes/No questions built from ten visual statements across 1,000 generated scenes and 46 labeled photographs, plus 1,000 object‑identification questions on the same photographs. The questions are designed to challenge reasoning with misleading local cues, nested logical operations, and both symbolic and natural‑language presentations, revealing significant gaps in current model performance.

By Yuzhou Wang, Emile Anand, Ijay Narang
arXiv AI
Aug 10

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.

By Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy