arXiv AI By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci

Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

Read the original on arXiv AI →

The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
4d ago

Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis

arXiv:2609.31456v1 Announce Type: new Abstract: Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hyp...

By Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers, Srinivasan Parthasarathy