arXiv AI By Jun Ling, Tao Huang, Junzhuo Liu, Bowen Tang, Peng Wang

GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

Read the original on arXiv AI →

arXiv:2607. 23913v1 Announce Type: new Abstract: Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream language-model inference cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 18

QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning

QCPruner is a training‑free visual token pruning method that conditions both token selection and representation on the query via bilateral utility weighting. It fuses keyword‑matched query anchors with cross‑modal cues to compute a nonnegative facility‑location objective that is monotone and submodular, guaranteeing a (1‑1/e) greedy approximation. Across multiple multimodal large language models, QCPruner consistently outperforms existing pruning methods, achieving over 96% of unpruned performance even with very few tokens retained.

By Shengli He (Guizhou University), Yongchao Liang (Guizhou University), Roumeng He (Shanghai Ocean University), Junjie Zeng (Guizhou University), Jiyuan He (Guizhou University), Can Wu (Guizhou University), Li Zheng (Guizhou University)
arXiv Computer Vision
3d ago

MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

arXiv:2609.34330v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visu...

By Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang, Weimin Ouyang, Rui Huang, Jiajun Cao, Sixiang Chen, Hao Jiang, Jixian Wu, Zheng Lu, Bofan Zhu, Renyuan Li, Shanghang Zhang