The paper introduces a training‑free visual token pruning strategy for vision‑language models that separates early vision‑guided pruning from later text‑guided reselection. By first pruning tokens with vision‑encoder attention, retaining candidates until the decoder midpoint, and then applying text‑to‑visual attention, the method preserves task‑relevant visual information. Across eight benchmarks and three models, it achieves an average performance recovery of 11.10 and 16.84 percentage points at 80% and 90% pruning, respectively, while maintaining comparable or lower LLM‑prefill latency.
arXiv:2605. 13178v2 Announce Type: replace-cross Abstract: In large vision-language models, visual tokens typically constitute the majority of input tokens, leading to substantial computational overhead.
By Sangin Lee, Yukyung Choi
ERASE is an adaptive two-stage token pruning framework designed to reduce computational overhead in Vision‑Language Models by eliminating redundant visual tokens. Stage 1 removes image‑level redundancy using lightweight raw‑image statistics, while Stage 2 progressively prunes instruction‑irrelevant tokens across decoder layers. Experiments on Qwen2.5‑VL‑7B show that ERASE retains 95.70 % of the original accuracy while keeping only 25 % of the tokens.
By Yuna Lee, Kyoungho Min, Yulhwa Kim
PACE introduces a training‑free Condense‑and‑Extract framework that speeds up Vision‑Language Model inference by first adaptively downsampling visual inputs before encoding and then selectively retaining essential tokens during decoding. The Adaptive Pixel Compressor (APC) reduces encoder workload while preserving global context, and the Dynamic Dual‑Attention Extractor (DDAE) keeps task‑critical details by fusing visual and language signals. Applied to Qwen2.5‑VL‑7B, PACE maintains 93.8% of performance using only 10% of visual tokens, achieving a 3.1× speedup in time to first token.
By Junjie Liu, Shengyuan Ye, Xu Chen
The paper introduces Adaptive Visual Token Pruning (AVTP), a training‑free framework that dynamically selects pruning layers and ratios for large vision‑language models (LVLMs) when processing multiple image sequences. By analyzing visual attention distributions across different LVLM architectures, AVTP adapts token retention to image importance, enabling efficient inference without relying on attention‑based computations incompatible with FlashAttention. Experiments show significant speedups—up to 2× for Qwen3VL‑8B—while preserving or even improving accuracy on multi‑image benchmarks.
By Rongyang Zhang, Chengqiang Lu, Cong Li, Hongchao Gu, Tingjia Shen, Xuyang Zhi, Qimeng Wang, Yan Gao, Yi Wu, Yao Hu, Hao Wang, Enhong Chen
LensVLM is an inference framework and post‑training recipe that lets Vision‑Language Models (VLMs) process compressed images of text by selectively expanding only the relevant parts back to full resolution. Using Qwen3.5‑9B‑Base, LensVLM achieves accuracy comparable to full‑text models at 4.3× compression and outperforms other compression baselines up to 10.1× across seven text QA benchmarks, while also improving performance on multimodal document and code tasks as compression increases.
By Roy Xie, Dan Friedman, Donghan Yu, Bowen Pan, Christopher Fifty, Jang-Hyun Kim, Xianzhi Du, Zhe Gan, Vivek Rathod, Bhuwan Dhingra
arXiv:2604.23950v2 Announce Type: replace
Abstract: Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in visual understanding and reasoning, but they also impose signif...
By Rinyoichi Takezoe, Yaqian Li, Zihao Bo, Anzhou Hou, Mo Guang, Kaiwen Long
arXiv:2609.24485v1 Announce Type: new
Abstract: Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction oft...
By Guangchuan Lv, Dianxing Shi, Dingjie FU
arXiv:2607.28627v2 Announce Type: replace-cross
Abstract: Long visual contexts challenge vision-language models: performance degrades as the number of distractors grows, and processing all tokens at...
By Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem
arXiv:2607. 28627v1 Announce Type: cross Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints.
By Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem
arXiv:2609.05916v1 Announce Type: cross
Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...
By Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We...