← Back to all news
arXiv Computer Vision September 23, 2026 By Yunge Li, Lanyu Xu

Neighbor-Aware Token Reduction via Hilbert Curve for Vision Transformers

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • efficiency
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 9

RAPID: Layer-Wise Redundancy-Aware Pruning and Importance-Driven Token Merging for Efficient ViT

arXiv:2606. 08156v1 Announce Type: cross Abstract: Vision Transformers (ViTs) achieve strong performance but suffer from high computational costs due to quadratic self-attention complexity.

By Kyumin Choi, Ikbeom Jang
llmsefficiency
More like this →
arXiv AI
Jul 10

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models

arXiv:2604. 11530v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have revolutionized multi-modal learning by jointly processing visual and textual information.

By Yvon Apedo, Martyna Poreba, Michal Szczepanski, Samia Bouchafa
llmsefficiencymultimodalsafety
More like this →
arXiv Computer Vision
Sep 22

VPRune: Efficient Training-free Pre-LLM Visual Token Pruning

arXiv:2609.24485v1 Announce Type: new Abstract: Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction oft...

By Guangchuan Lv, Dianxing Shi, Dingjie FU
llmsefficiencymultimodalbenchmarkssafety
More like this →
arXiv Machine Learning
Aug 21

Clustering and Token Denoising for Faster and More Robust VLMs

arXiv:2608. 19285v1 Announce Type: cross Abstract: Recent Visual-Language Models (VLMs) have enhanced the capabilities of pre-trained LLMs by adding vision tokens alongside text, with approaches like LLaVA showing impressive results.

By Baptiste Rossigneux, Inna Kucher, Vincent Lorrain, Emmanuel Casseau
llmsefficiencymultimodalbenchmarks
More like this →
arXiv AI
Sep 1

Centering before Pruning: Lightweight Geometry Correction for Diversity-Based Visual Token Pruning in LVLMs

arXiv:2608.30263v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) incur substantial inference costs due to their long and highly redundant visual-token sequences. Diversity-based...

By Shunjie Wen, Jaeyeon Lee, Dong-Wan Choi
llmsefficiencymultimodalbenchmarks
More like this →
arXiv AI
Jun 2

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

arXiv:2605. 13178v2 Announce Type: replace-cross Abstract: In large vision-language models, visual tokens typically constitute the majority of input tokens, leading to substantial computational overhead.

By Sangin Lee, Yukyung Choi
llmsfine-tuningefficiencymultimodal
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea