arXiv AI

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics

arXiv:2606. 03569v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated remarkable capabilities but suffer from significant computational overhead during inference.

arXiv Computer Vision
Aug 27

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

The paper shows that only a small subset of attention heads in vision-language models is responsible for selecting critical visual tokens. By pruning tokens based on similarity before LLM reasoning and then applying head‑aware pruning during reasoning, the proposed ProViP framework achieves high task performance with significant speedups. Experiments on LLaVA‑1.5‑7B demonstrate 95.9% performance retention and a 1.62× inference speedup at an 88.9% pruning ratio.

By Chaofang Ma, Lin Jiang, Carol Jingyi Li, Xingyu Liu, Zeyu Li, Jiang Xu, Wei Zhang
Hugging Face Trending Papers
5d ago

When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

The paper introduces a training‑free visual token pruning strategy for vision‑language models that separates early vision‑guided pruning from later text‑guided reselection. By first pruning tokens with vision‑encoder attention, retaining candidates until the decoder midpoint, and then applying text‑to‑visual attention, the method preserves task‑relevant visual information. Across eight benchmarks and three models, it achieves an average performance recovery of 11.10 and 16.84 percentage points at 80% and 90% pruning, respectively, while maintaining comparable or lower LLM‑prefill latency.

arXiv AI
Sep 10

STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

arXiv:2609.05916v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...

By Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang
arXiv Computer Vision
Sep 23

Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs

The paper introduces STD, a hierarchical token pruning framework for Large Vision‑Language Models that aligns pruning strategies with the functional roles of different network stages. By using high‑frequency spectral analysis in shallow layers, Gaussian‑smoothed attention in intermediate layers, and a stability‑adaptive trigger in deep layers, STD preserves essential visual information while aggressively reducing token counts. Experiments demonstrate that STD outperforms existing pruning methods, achieving up to 94.4% token reduction and a 3.9× speed‑up on LLaVA‑NeXT‑7B.

By Shuo Zhang, Jintao Tong, Yixiong Zou, Yuhua Li, Ruixuan Li
arXiv AI
Jun 26

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference

arXiv:2606. 27161v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the large number of visual tokens, which introduces substantial computational overhead.

By Tinghao Wang, Yichen Guo, Rui Huang, Zheng Lu, Qizhe Zhang, Chenxi Li, Yuan Zhang, Jiajun Cao, Zhirong Shen, Yaosong Du, Guangyan Gan, Wenya Wang, Lin William Cong, Shanghang Zhang