arXiv:2606. 19489v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) enhance interpretability by projecting learned features into a human-understandable concept space.
By Ya Wang, Adrian Paschke
The paper introduces CFM, a language‑aligned concept foundation model for vision that generates fine‑grained, human‑interpretable concepts with spatial grounding. By pairing CFM with a strong semantic foundation model, it provides explanations for downstream tasks such as classification, segmentation, and captioning. The authors also analyze local co‑occurrence of concepts to define relationships, improving concept naming and yielding richer explanations while maintaining competitive performance.
By Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi, Bernt Schiele, Jonas Fischer
arXiv:2605. 18160v2 Announce Type: replace-cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information.
By Xinpeng Dong, Min Zhang, Kairong Han, Xu Tan, Fei Wu, Kun Kuang
arXiv:2606. 19882v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) enhance the interpretability of deep learning networks by aligning the features extracted from images with natural concepts.
By Tongqing Shi, Ge Yan, Tuomas Oikarinen, Tsui-Wei Weng
arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.
By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
arXiv:2606. 30498v1 Announce Type: cross Abstract: Human decision-making interprets the world through high-level concepts, such as recognizing a bird by its belly color.
By Laines Schmalwasser, Jan Blunk, Niklas Penzel, Julia Niebling, Joachim Denzler
Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can...
arXiv:2609.23717v1 Announce Type: new
Abstract: Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which regi...
By Liuyang Song, Yi Zhang, Zhongyi Deng, Daqian Yang, Hongbo Zhang
arXiv:2607. 21371v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories.
By Sung-Hoon Yoon, Hoyong Kwon, Changgyoon Oh, Kuk-Jin Yoon
arXiv:2608.22584v1 Announce Type: new
Abstract: Two-stage neuro-symbolic architectures provide an elegant paradigm for visual problem solving by cleanly separating connectionist perception of predefi...
By Sparsh Tiwari, Gesina Schwalbe, Bettina Finzel
arXiv:2608. 07886v1 Announce Type: cross Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region.
By Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna
arXiv:2603. 26798v2 Announce Type: replace-cross Abstract: Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this space is rarely inspected.
By Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika M\"utze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann