arXiv:2608.29313v1 Announce Type: cross
Abstract: CLIP-like vision-language models (VLMs) trained with contrastive objectives learn strong global image-text representations, but their Euclidean embed...
By Matin Mahmood, Antonio Rueda-Toicen, Mohamed ElBassat, Seifeldin Elkerdany, Weixing Wang, Gerard de Melo
arXiv:2603. 26798v2 Announce Type: replace-cross Abstract: Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this space is rarely inspected.
By Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika M\"utze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann
arXiv:2607. 03143v1 Announce Type: cross Abstract: Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details.
By Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu
arXiv:2606. 29464v1 Announce Type: cross Abstract: Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets.
By Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon
arXiv:2606.23843v2 Announce Type: replace
Abstract: Vision-language models (VLMs) achieve strong cross-modal alignment but remain brittle to negation, often relying on shallow word associations rathe...
By Hoang-Bao Le, Aiden Durrant, Thai Son Mai, Binh T. Nguyen, Liting Zhou, Cathal Gurrin
arXiv:2609.24276v1 Announce Type: new
Abstract: Hyperbolic vision-language models (VLMs) represent image and text features in a geometry naturally suited to hierarchy, but their adaptation to downstr...
By Andro Erdelez, Pascal Mettes, Behzad Bozorgtabar
SOCO is a new benchmark for Semantic Object Correspondence that introduces a taxonomy of correspondence types and provides consistent, functionally meaningful keypoint annotations across 100 categories and over 1M correspondence pairs. It also includes keypoint language descriptions, enabling evaluation of large vision‑language models and their fine‑grained part‑level understanding. Experiments show that vision foundation backbones encode strong semantic structure but transfer correspondences poorly across related categories, LVLMs excel at text‑prompted part localization but lag in visual‑reference matching, and correspondence performance predicts dense downstream tasks more strongly than ImageNet classification.
By Olaf D\"unkel, Basavaraj Sunagad, Haoran Wang, David T. Hoffmann, Christian Theobalt, Adam Kortylewski
arXiv:2609.24564v1 Announce Type: new
Abstract: CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encode...
By Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen
arXiv:2607. 02909v1 Announce Type: cross Abstract: Taxonomies provide key information about the semantic relationships between concepts and the inherent organization of vision and language.
By Hulingxiao He, Zhi Tan, Yuxin Peng
arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.
By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
The paper introduces CFM, a language‑aligned concept foundation model for vision that generates fine‑grained, human‑interpretable concepts with spatial grounding. By pairing CFM with a strong semantic foundation model, it provides explanations for downstream tasks such as classification, segmentation, and captioning. The authors also analyze local co‑occurrence of concepts to define relationships, improving concept naming and yielding richer explanations while maintaining competitive performance.
By Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi, Bernt Schiele, Jonas Fischer
The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci