arXiv AI

Formalizing the Binding Problem

arXiv:2606. 03976v1 Announce Type: cross Abstract: Representations of the world, arguably, contain information about features (e.

arXiv AI
6d ago

Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.

By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv AI
Aug 12

Token-Based Detection of Spurious Correlations in Vision Transformers

arXiv:2509. 04009v2 Announce Type: replace-cross Abstract: Due to their powerful feature association capabilities, neural network-based computer vision models have the ability to detect and exploit unintended patterns within the data, potentially leading to correct predictions based on incorrect or unintended but statistically relevant signals.

By Solha Kang, Esla Timothy Anzaku, Wesley De Neve, Arnout Van Messem, Joris Vankerschaver, Francois Rameau, Utku Ozbulak
arXiv AI
Sep 7

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

The paper surveys how commonsense reasoning is being integrated into computer vision, moving beyond traditional CNNs that only detect objects. It reviews methods that use knowledge graphs, scene graphs, neuro-symbolic models, and transformers to add contextual understanding, thereby improving object recognition and spatial reasoning. The authors also discuss current limitations such as dataset bias and knowledge gaps, and propose future research directions in cross‑modal reasoning, scalable knowledge injection, and hybrid architectures.

By Bahar Uddin Mahmud, Sumit Barua, Guan Yue Hong, Ajay Gupta, Hexu Liu