The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv:2603. 22278v2 Announce Type: replace-cross Abstract: Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations.
By Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina, David Bau, Antonio Torralba, Tamar Rott Shaham
arXiv:2606. 13288v1 Announce Type: cross Abstract: Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding.
By Wei Li, Zhen Huang, Xinmei Tian
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.
By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
arXiv:2609.23717v1 Announce Type: new
Abstract: Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which regi...
By Liuyang Song, Yi Zhang, Zhongyi Deng, Daqian Yang, Hongbo Zhang
arXiv:2509. 04009v2 Announce Type: replace-cross Abstract: Due to their powerful feature association capabilities, neural network-based computer vision models have the ability to detect and exploit unintended patterns within the data, potentially leading to correct predictions based on incorrect or unintended but statistically relevant signals.
By Solha Kang, Esla Timothy Anzaku, Wesley De Neve, Arnout Van Messem, Joris Vankerschaver, Francois Rameau, Utku Ozbulak