arXiv Computer Vision By Zesheng Li, Chengchang Pan, Honggang Qi

GraSP-VL: Length as a Semantic Granularity Interface for Vision-Language Representations

Read the original on arXiv Computer Vision →

GraSP‑VL demonstrates that the length of frozen vision‑language embeddings can serve as a controllable semantic interface. By learning a shared near‑orthogonal prefix transform, the method creates a Semantic Matryoshka where short prefixes encode coarse semantics and longer prefixes reveal finer language‑grounded distinctions, all while preserving the original embedding geometry. Experiments on COCO/Flickr30K and SugarCrepe‑clean show strong performance with negligible drift in the full embedding space.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 24

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.

By Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv AI
6d ago

Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.

By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci