arXiv Computer Vision By Shengli He, Yongchao Liang, Roumeng He, Junjie Zeng, Jiyuan He, Xin Fang, Can Wu, Li Zheng

From Token Importance to Conditional Removability: Rethinking Visual Token Pruning in Multimodal Large Language Models

Read the original on arXiv Computer Vision →

The paper argues that token importance alone is insufficient to determine safe removal of visual tokens in multimodal large language models, because removability depends on representation depth and the surrounding deletion set. Through controlled experiments, the authors show that the same tokens can have different effects when removed at different depths or contexts. They introduce CoRePrune, a training‑free two‑stage pruning framework that refreshes deletion effects as visual representations evolve and refines candidate tokens based on the current deletion set, achieving high performance retention across multiple backbones and reducing prefill time significantly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 2

SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

SinkPruner is a training‑free framework that prunes visual tokens for multimodal large language models by first removing high‑norm redundant tokens with a visual sanitizer and then selectively keeping tokens that align with the text query using a text‑guided pruner. The coarse‑to‑fine design reduces attention sink and dispersion, enabling an 89% token reduction while preserving 96.5% of LLaVA‑1.5’s performance and 91.8% of Qwen2.5‑VL’s performance across twelve image‑language and four video‑language benchmarks. The visual sanitizer also improves existing pruning methods, showing strong transferability.

By Shiyu Li, Zi-Yuan Hu, Shijia Huang, Yanyang Li, Yiwu Zhong, Liwei Wang
arXiv AI
Sep 10

STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

arXiv:2609.05916v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...

By Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang
arXiv Computer Vision
23h ago

MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

arXiv:2609.34330v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visu...

By Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang, Weimin Ouyang, Rui Huang, Jiajun Cao, Sixiang Chen, Hao Jiang, Jixian Wu, Zheng Lu, Bofan Zhu, Renyuan Li, Shanghang Zhang
arXiv Computer Vision
Aug 27

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

The paper shows that only a small subset of attention heads in vision-language models is responsible for selecting critical visual tokens. By pruning tokens based on similarity before LLM reasoning and then applying head‑aware pruning during reasoning, the proposed ProViP framework achieves high task performance with significant speedups. Experiments on LLaVA‑1.5‑7B demonstrate 95.9% performance retention and a 1.62× inference speedup at an 88.9% pruning ratio.

By Chaofang Ma, Lin Jiang, Carol Jingyi Li, Xingyu Liu, Zeyu Li, Jiang Xu, Wei Zhang