arXiv AI By Junchi Liao, Jiawen Deng, Fuji Ren

VISTA: Auditing Semantic Divergence in Vision-Language Models

Read the original on arXiv AI →

arXiv:2607. 02995v1 Announce Type: cross Abstract: Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideological symbols, some models produce unusually uniform responses that differ from what peer models say about the same input.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 21

DiaVLo: Diagnosing Behaviours of Vision-Language Models

DiaVLo is a diagnostic framework for vision‑language models (VLMs) that uses human curation and VLM generation to create specifications of desired and observed behaviours, revealing potential misalignments. It also offers causal estimates to pinpoint the most influential concepts driving VLM behaviour. Experiments on several open‑source VLMs under classification and generation tasks show that DiaVLo’s behaviour labels correlate with model performance and illuminate how VLMs perceive, organise, and prioritise concepts.

By Lorenzo Corti, Jie Yang
arXiv Computation and Language
Sep 2

Reliability Challenges in Diffusion Vision-Language Models

The paper presents the first systematic reliability evaluation of diffusion-based Large Vision‑Language Models (dLVLMs), comparing six diffusion models to autoregressive (AR) baselines across four dimensions. Key findings include a reversal of the yes‑bias seen in AR models for binary visual queries, competitive hallucination rates but lower linguistic quality, near‑zero accuracy for underrepresented racial groups with opposite‑polarity gender bias, and accuracy collapse in multiple‑choice tasks when the correct option is shorter than distractors due to a length prior emerging at the first denoising step. Additionally, tokens committed late in denoising with low confidence correlate with hallucinated content, indicating a unique mechanistic signal in diffusion generation.

By Md. Atabuzzaman, Chris Thomas
arXiv AI
Sep 16

Same Answer, Different Representations: Hidden instability in VLMs

arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...

By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini