arXiv Machine Learning By Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh, Nikolai Rozanov, Kentaro Inui

The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs

Read the original on arXiv Machine Learning →

arXiv:2607. 09544v1 Announce Type: cross Abstract: Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 16

Same Answer, Different Representations: Hidden instability in VLMs

arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...

By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini
arXiv AI
Sep 21

DiaVLo: Diagnosing Behaviours of Vision-Language Models

DiaVLo is a diagnostic framework for vision‑language models (VLMs) that uses human curation and VLM generation to create specifications of desired and observed behaviours, revealing potential misalignments. It also offers causal estimates to pinpoint the most influential concepts driving VLM behaviour. Experiments on several open‑source VLMs under classification and generation tasks show that DiaVLo’s behaviour labels correlate with model performance and illuminate how VLMs perceive, organise, and prioritise concepts.

By Lorenzo Corti, Jie Yang