V-CoLA: Vision Token Compression with Linear Attention
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2604.23950v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in visual understanding and reasoning, but they also impose signif...
The paper shows that only a small subset of attention heads in vision-language models is responsible for selecting critical visual tokens. By pruning tokens based on similarity before LLM reasoning and then applying head‑aware pruning during reasoning, the proposed ProViP framework achieves high task performance with significant speedups. Experiments on LLaVA‑1.5‑7B demonstrate 95.9% performance retention and a 1.62× inference speedup at an 88.9% pruning ratio.
The paper introduces Adaptive Visual Token Pruning (AVTP), a training‑free framework that dynamically selects pruning layers and ratios for large vision‑language models (LVLMs) when processing multiple image sequences. By analyzing visual attention distributions across different LVLM architectures, AVTP adapts token retention to image importance, enabling efficient inference without relying on attention‑based computations incompatible with FlashAttention. Experiments show significant speedups—up to 2× for Qwen3VL‑8B—while preserving or even improving accuracy on multi‑image benchmarks.
arXiv:2604. 00757v2 Announce Type: replace-cross Abstract: Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens.
arXiv:2609.05916v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...
VETO (Video Efficient Token Optimization for Vision Language Models) is a plug‑in that reduces the quadratic cost of visual tokens in long‑video inference by applying dual‑axis compression: an intra‑frame compressor merges semantically similar tokens within each frame, and an inter‑frame compressor merges temporally redundant frames. By first compressing spatial dimensions, VETO lowers the cost of subsequent global temporal matching, surpassing single‑axis methods and achieving up to 45% faster inference on models such as LLaVA‑OneVision‑7B while maintaining or improving accuracy. The approach is universally applicable across LLaVA‑OneVision, InternVL‑2.5, and LongVA, preserving or enhancing zero‑shot accuracy even under extreme token budgets.