arXiv Computer Vision By Yuna Lee, Kyoungho Min, Yulhwa Kim

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning

Read the original on arXiv Computer Vision →

ERASE is an adaptive two-stage token pruning framework designed to reduce computational overhead in Vision‑Language Models by eliminating redundant visual tokens. Stage 1 removes image‑level redundancy using lightweight raw‑image statistics, while Stage 2 progressively prunes instruction‑irrelevant tokens across decoder layers. Experiments on Qwen2.5‑VL‑7B show that ERASE retains 95.70 % of the original accuracy while keeping only 25 % of the tokens.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 28

Multi-Image Visual Token Pruning in Large Visual Language Models

The paper introduces Adaptive Visual Token Pruning (AVTP), a training‑free framework that dynamically selects pruning layers and ratios for large vision‑language models (LVLMs) when processing multiple image sequences. By analyzing visual attention distributions across different LVLM architectures, AVTP adapts token retention to image importance, enabling efficient inference without relying on attention‑based computations incompatible with FlashAttention. Experiments show significant speedups—up to 2× for Qwen3VL‑8B—while preserving or even improving accuracy on multi‑image benchmarks.

By Rongyang Zhang, Chengqiang Lu, Cong Li, Hongchao Gu, Tingjia Shen, Xuyang Zhi, Qimeng Wang, Yan Gao, Yi Wu, Yao Hu, Hao Wang, Enhong Chen
arXiv AI
Aug 5

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models

arXiv:2608. 03112v1 Announce Type: cross Abstract: Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deployment on resource-constrained edge devices and in real-time surveillance applications.

By Paribesh Regmi, Qingshuang Chen, Chi Zhang, Heba Aly, Yelin Kim, Hongda Mao
arXiv Computer Vision
3d ago

Adaptive Visual Token Reduction for Accelerated Image Understanding

The paper introduces ReFIT, an instruction‑guided visual token reduction framework designed to accelerate large vision‑language model inference. ReFIT combines Relevance‑Guided Window Reshaping (RWR) to adaptively capture instruction‑relevant regions and Instruction‑Guided Token Refinement (ITR) to prune unnecessary visual tokens. Experiments on four VQA benchmarks show that ReFIT improves answer accuracy while lowering computational cost, and qualitative results confirm its ability to localize relevant regions and remove extraneous visual information.

By Seyoung Jeong, Jong Pil Yun, Sang Jun Lee
arXiv Computer Vision
Sep 23

From Token Importance to Conditional Removability: Rethinking Visual Token Pruning in Multimodal Large Language Models

The paper argues that token importance alone is insufficient to determine safe removal of visual tokens in multimodal large language models, because removability depends on representation depth and the surrounding deletion set. Through controlled experiments, the authors show that the same tokens can have different effects when removed at different depths or contexts. They introduce CoRePrune, a training‑free two‑stage pruning framework that refreshes deletion effects as visual representations evolve and refines candidate tokens based on the current deletion set, achieving high performance retention across multiple backbones and reducing prefill time significantly.

By Shengli He, Yongchao Liang, Roumeng He, Junjie Zeng, Jiyuan He, Xin Fang, Can Wu, Li Zheng