arXiv Computer Vision

Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores

arXiv Machine Learning
Jul 2

Information-Regularized Attention for Visual-Centric Reasoning

arXiv:2607. 00434v1 Announce Type: cross Abstract: Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting after full-parameter instruction tuning.

By Guohao Sun, Xiaofang Wang, Yash Patel, Mengchen Liu, Zhiqiang Tao, Praveen Krishnan
arXiv Computer Vision
Sep 25

IronViT: Toward Efficient Generalist Visual Representation Learning

IronViT proposes a new approach to building efficient generalist vision encoders by first consolidating the knowledge of multiple specialist teachers into a softmax attention bridge and then transferring this consolidated representation to a hybrid softmax‑linear attention architecture. This two‑stage distillation process, supported by a curated data pipeline, allows the model to capture semantic, spatial, language‑aligned, and action‑relevant cues while avoiding the high‑resolution cost of traditional softmax attention. Across tasks such as recognition, retrieval, dense prediction, multimodal understanding, and robotic learning, IronViT matches or exceeds the performance of leading specialist and generalist encoders, with the hybrid encoder offering increasing efficiency at higher resolutions.

By Jiaxi Huang, Yueqi Hu, Xin Zhu, Xiaopeng Zhang, Huiting Qiao, Yanglin Zhang, Zefeng Ji, Rongxue Li, Yifei Xu, Huiying Yu, Wei Liu, Jiayin Zheng, Yinggan Xu, Peipeng Chen, Yin Zhang, Jian Yao
arXiv Machine Learning
Aug 27

Token-Oriented Semantic Communication with Pretrained Vision Transformers

The paper introduces a token‑oriented semantic communication framework that transmits only task‑relevant image latents instead of full token embeddings, reducing communication cost and improving interoperability. It leverages a spatial alignment between vision transformer patch tokens and learned image compression latents, enabling token‑level relevance estimation and selective transmission. Experiments on ImageNet demonstrate a superior rate–accuracy trade‑off compared to existing semantic communication methods and hand‑crafted codecs.

By Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim
arXiv Computer Vision
Sep 24

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.

By Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv AI
Sep 10

STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

arXiv:2609.05916v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...

By Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang