Hugging Face Trending Papers

ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

Hugging Face Trending Papers
Sep 28

When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

The paper introduces a training‑free visual token pruning strategy for vision‑language models that separates early vision‑guided pruning from later text‑guided reselection. By first pruning tokens with vision‑encoder attention, retaining candidates until the decoder midpoint, and then applying text‑to‑visual attention, the method preserves task‑relevant visual information. Across eight benchmarks and three models, it achieves an average performance recovery of 11.10 and 16.84 percentage points at 80% and 90% pruning, respectively, while maintaining comparable or lower LLM‑prefill latency.

arXiv Computer Vision
6d ago

Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs

Gaze Attention is a new mechanism for multimodal large language models that selectively focuses on relevant visual regions during each generation step, rather than attending to all visual tokens. By grouping tokens into spatial regions and using learnable context tokens to retain global information, it reduces attention computation and visual key‑value entries by up to 90%. Experiments on 13 image and 6 video benchmarks show that Gaze Attention matches or outperforms dense‑attention baselines while using fewer visual resources.

By Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun
arXiv Computer Vision
Aug 31

Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation

The paper introduces QK Product Steering, a data‑free, training‑free method that edits the query‑key product in vision‑language models to reduce object hallucination. By suppressing a few dominant singular modes in selected middle layers and mapping the edited product back to query weights, the approach lowers hallucination rates without affecting inference cost. Experiments on three GQA‑based VLMs show a 4.0% average reduction in CHAIR$_s$, with the effect localized to symmetric mutual‑attention channels.

By Karn Tiwari, Varnith Chordia, Prathosh A P
arXiv Computer Vision
Sep 23

Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs

The paper introduces STD, a hierarchical token pruning framework for Large Vision‑Language Models that aligns pruning strategies with the functional roles of different network stages. By using high‑frequency spectral analysis in shallow layers, Gaussian‑smoothed attention in intermediate layers, and a stability‑adaptive trigger in deep layers, STD preserves essential visual information while aggressively reducing token counts. Experiments demonstrate that STD outperforms existing pruning methods, achieving up to 94.4% token reduction and a 3.9× speed‑up on LLaVA‑NeXT‑7B.

By Shuo Zhang, Jintao Tong, Yixiong Zou, Yuhua Li, Ruixuan Li