arXiv AI

DiaVLo: Diagnosing Behaviours of Vision-Language Models

DiaVLo is a diagnostic framework for vision‑language models (VLMs) that uses human curation and VLM generation to create specifications of desired and observed behaviours, revealing potential misalignments. It also offers causal estimates to pinpoint the most influential concepts driving VLM behaviour. Experiments on several open‑source VLMs under classification and generation tasks show that DiaVLo’s behaviour labels correlate with model performance and illuminate how VLMs perceive, organise, and prioritise concepts.

arXiv AI
Aug 18

OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models

arXiv:2602. 18094v2 Announce Type: replace-cross Abstract: Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and identically distributed (IID).

By Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang, Yang Zhou, Huazhu Fu, Jingrun Chen
arXiv Computer Vision
Sep 22

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

arXiv:2609.24576v1 Announce Type: cross Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...

By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
arXiv AI
Sep 16

Same Answer, Different Representations: Hidden instability in VLMs

arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...

By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini
arXiv Computer Vision
Sep 11

Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World Conditions

The paper investigates using Vision Language Models (VLMs) to accelerate verification and validation (V&V) of classification models by automatically detecting systematic errors. It introduces a VLM-based error slice detection (ESD) method that groups and labels errors, demonstrating its ability to identify perturbations in a non-military dataset and to cluster images by surroundings in a military context. The study highlights challenges such as underrepresentation of defence data in VLM training and limited contextual diversity, and suggests that while fully automated V&V is not yet feasible, VLMs could speed up the process in the future.

By Dieuwertje Alblas, Alma M. Liezenga, Jan Erik van Woerden, Fedor Taggenbrock, Dalia Aljawaheri, Klamer Schutte
arXiv Computation and Language
Sep 4

FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models

FPCO-Dialog is a new benchmark designed to evaluate how vision‑language models correct and cooperate when faced with repeated false premises in multi‑turn dialogues. The dataset contains 1,080 images and 10,800 question turns, organized by visual complexity, object category, and false‑premise class, and follows a 10‑turn protocol where a correct dialogue prefix is followed by repeated false‑premise expressions. Using a model‑agnostic protocol and the CorrTP@K correction‑rate metric, the benchmark reveals significant differences among 20 commercial and open‑source VLMs in their correction tendencies, turn‑wise dynamics, and responses to different false‑premise types.

By Jiayuan Ma, Yuqi Lu, Weiyang Guo, Chenrui Wang, Junyi Shu, Xuebo Liu, Min Zhang, Jing Li