arXiv:2608.29313v1 Announce Type: cross
Abstract: CLIP-like vision-language models (VLMs) trained with contrastive objectives learn strong global image-text representations, but their Euclidean embed...
By Matin Mahmood, Antonio Rueda-Toicen, Mohamed ElBassat, Seifeldin Elkerdany, Weixing Wang, Gerard de Melo
arXiv:2607. 02909v1 Announce Type: cross Abstract: Taxonomies provide key information about the semantic relationships between concepts and the inherent organization of vision and language.
By Hulingxiao He, Zhi Tan, Yuxin Peng
arXiv:2609.24564v1 Announce Type: new
Abstract: CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encode...
By Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen
Prompt learning modifies vision‑language models by optimizing continuous prompt vectors, yet the resulting prompts are hard to interpret in natural language. PromptSpLiCE is a post‑hoc method that rewrites each class‑conditioned text embedding as a sparse mix of concepts from a fixed dictionary, enabling a direct comparison of concept profiles before and after prompt learning. Across 11 image‑classification datasets, the method shows that only about 1.6 of the initial top‑10 concepts remain after learning, and that larger profile changes correlate with higher accuracy gains, while a derived gradient expression offers geometric insight into loss sensitivity.
By Ryo Kamiya, Hiroshi Kera, Kazuhiko Kawamoto
arXiv:2607. 00684v1 Announce Type: new Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts.
By Seokhee Jin, Changhwan Sung, Sunung Mun, Hoyoung Kim, Jungseul Ok
arXiv:2608.21819v1 Announce Type: cross
Abstract: Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions wh...
By Jihyung Ko, Eunji Jung, Hyeongsub Kim, Ziseok Lee, Jae Won Cho, Sanghyun Jo, Kyungsu Kim
arXiv:2606. 00275v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have demonstrated impressive performance on multimodal tasks through scaled architectures and extensive training.
By Zijie Zhou, Dandan Zhu, Hangxiangpan Wang, Heng Zhang, Huishen Jiao, Yi Zhao
arXiv:2606.23843v2 Announce Type: replace
Abstract: Vision-language models (VLMs) achieve strong cross-modal alignment but remain brittle to negation, often relying on shallow word associations rathe...
By Hoang-Bao Le, Aiden Durrant, Thai Son Mai, Binh T. Nguyen, Liting Zhou, Cathal Gurrin
arXiv:2603. 22042v3 Announce Type: replace-cross Abstract: While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part-to-whole or parent-child structures, and often face challenges in multi-object compositional scenarios.
By Hayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young Chun
arXiv:2606. 16193v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret.
By Yusong Zhao, Hengyi Wang, Tanuja Ganu, Akshay Nambi, Hao Wang
arXiv:2603. 09493v2 Announce Type: replace-cross Abstract: The adaptation of large-scale vision-language models (VLMs) to downstream tasks with limited labeled data remains a significant challenge.
By Enming Zhang, Jiayang Li, Yanlong Wang, Yanru Wu, Zhenyu Liu, Yang Li
The paper introduces EvoPrompt, a framework for adapting vision‑language models to new tasks with limited data while preventing catastrophic forgetting. EvoPrompt uses a Modality‑Shared Prompt Projector to create hierarchical prompts and an evolutionary training strategy that separates low‑rank updates into directional and magnitude components, preserving learned semantic directions. Experiments show that EvoPrompt achieves state‑of‑the‑art few‑shot performance while maintaining the original zero‑shot capabilities of the pre‑trained models.
By Enming Zhang, Jiayang Li, Yanlong Wang, Yanru Wu, Zhenyu Liu, Yang Li