arXiv:2609.24564v1 Announce Type: new
Abstract: CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encode...
By Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen
arXiv:2609.01147v1 Announce Type: cross
Abstract: Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders strug...
By Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang, Yu Rong, Hong Cheng, Hou Pong Chan, Chenghao Xiao
GraSP‑VL demonstrates that the length of frozen vision‑language embeddings can serve as a controllable semantic interface. By learning a shared near‑orthogonal prefix transform, the method creates a Semantic Matryoshka where short prefixes encode coarse semantics and longer prefixes reveal finer language‑grounded distinctions, all while preserving the original embedding geometry. Experiments on COCO/Flickr30K and SugarCrepe‑clean show strong performance with negligible drift in the full embedding space.
By Zesheng Li, Chengchang Pan, Honggang Qi
arXiv:2607. 07047v1 Announce Type: cross Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety.
By Szczepan Konior, Alexandre Quemy, Przemys{\l}aw Klocek, Gr\'egoire Cattan, Bart{\l}omiej Sobieski
arXiv:2609.37230v1 Announce Type: cross
Abstract: Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a disti...
By Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim
arXiv:2607. 20092v1 Announce Type: cross Abstract: Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful.
By Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani