Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on similarity-based guidance, which exploits pairwise text-vision and vision-vision token correlations for compression.
arXiv:2609.24485v1 Announce Type: new
Abstract: Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction oft...
By Guangchuan Lv, Dianxing Shi, Dingjie FU
The paper introduces a training‑free visual token pruning strategy for vision‑language models that separates early vision‑guided pruning from later text‑guided reselection. By first pruning tokens with vision‑encoder attention, retaining candidates until the decoder midpoint, and then applying text‑to‑visual attention, the method preserves task‑relevant visual information. Across eight benchmarks and three models, it achieves an average performance recovery of 11.10 and 16.84 percentage points at 80% and 90% pruning, respectively, while maintaining comparable or lower LLM‑prefill latency.
The paper introduces Adaptive Visual Token Pruning (AVTP), a training‑free framework that dynamically selects pruning layers and ratios for large vision‑language models (LVLMs) when processing multiple image sequences. By analyzing visual attention distributions across different LVLM architectures, AVTP adapts token retention to image importance, enabling efficient inference without relying on attention‑based computations incompatible with FlashAttention. Experiments show significant speedups—up to 2× for Qwen3VL‑8B—while preserving or even improving accuracy on multi‑image benchmarks.
By Rongyang Zhang, Chengqiang Lu, Cong Li, Hongchao Gu, Tingjia Shen, Xuyang Zhi, Qimeng Wang, Yan Gao, Yi Wu, Yao Hu, Hao Wang, Enhong Chen
arXiv:2609.23715v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) incur substantial computational overhead due to the reliance on hundreds of visual tokens to represent images....
By Yahong Wang, Zhangkai Ni, Juncheng Wu, Yuyin Zhou, Ying Wen, Lianghua He
SinkPruner is a training‑free framework that prunes visual tokens for multimodal large language models by first removing high‑norm redundant tokens with a visual sanitizer and then selectively keeping tokens that align with the text query using a text‑guided pruner. The coarse‑to‑fine design reduces attention sink and dispersion, enabling an 89% token reduction while preserving 96.5% of LLaVA‑1.5’s performance and 91.8% of Qwen2.5‑VL’s performance across twelve image‑language and four video‑language benchmarks. The visual sanitizer also improves existing pruning methods, showing strong transferability.
By Shiyu Li, Zi-Yuan Hu, Shijia Huang, Yanyang Li, Yiwu Zhong, Liwei Wang