arXiv:2511. 16527v2 Announce Type: replace-cross Abstract: Contrastive vision-language models continue to be the dominant approach for image-text retrieval.
By Kwun Ho Ngan, Saman Sadeghi Afgeh, Joe Townsend, Artur d'Avila Garcez
arXiv:2609.05730v1 Announce Type: cross
Abstract: Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation...
By Samir Char, Carles Domingo-Enrich, Randall Balestriero
arXiv:2606. 16799v1 Announce Type: cross Abstract: Existing vision-language model (VLM)-based AI-generated image quality assessment (AIGIQA) methods suffer from a fundamental semantic-distortion dimensional conflict: monolithic representations optimized for semantic discrimination inherently entangle compositional understanding with low-level perceptual sensitivity, rendering them blind to fine-grained quality degradations.
By Zijie Meng
arXiv:2609.07937v1 Announce Type: cross
Abstract: Structured visual reasoning, such as image puzzles, demands fine-grained visual perception, an ability current Vision Language Models (VLMs) lack. VL...
By Harsha Patnala, Debopriyo Banerjee, Ayush Sunil Munot, Somak Aditya
arXiv:2609.10224v1 Announce Type: new
Abstract: Vision-language models such as CLIP embed images and text in a shared space, where modality-specific distributions often remain separated. Existing acc...
By Zonglin Yang, Huilan Ma, Xudan Zheng, Yuejun Xie
arXiv:2606.23843v2 Announce Type: replace
Abstract: Vision-language models (VLMs) achieve strong cross-modal alignment but remain brittle to negation, often relying on shallow word associations rathe...
By Hoang-Bao Le, Aiden Durrant, Thai Son Mai, Binh T. Nguyen, Liting Zhou, Cathal Gurrin