The paper introduces EviSpec, a training‑free compiler that generates complementary evidence specifications to improve high‑resolution multimodal large language models (MLLMs). By explicitly guiding visual search with structured evidence specifications, EviSpec achieves significant relative gains—up to 14.8% over random evidence—across five MLLMs and three benchmarks, and also sets new state‑of‑the‑art results on VQA and hallucination‑focused tasks.
By Zhongkuan Mao, Wenzhuo Zhao, Xianjie Liu, Yidong Wang, Zhao Gao, Ronghao Xian, Yao Jiang, Yi Zhang, Liangjian Wen, Keren Fu
VisCache introduces a two-stage, plug‑and‑play framework for pruning visual key‑value caches in Vision Large Language Models without retraining. The first stage filters out temporally redundant keyframes, while the second stage, PruneKV, applies a parabolic layer‑wise budget and asymmetric update to selectively prune keys and fuse values, preserving essential context. Experiments show up to 2.35× speedup and significant memory savings with only 19–28% of the original cache retained, outperforming existing baselines.
By Lyuke Wang, Zhuo Li, Guangxu Zhu
FAVE (Foveated Adaptive Visual Encoding) is a lightweight, variable‑resolution Vision Transformer that encodes user‑selected image regions at high acuity while maintaining the image’s native geometry. In controlled experiments on small‑object ImageNet crops, FAVE outperforms a fixed‑resolution ViT by 9.4 top‑1 points while using 12.7× fewer FLOPs. When added as a local branch to FastVLM, FAVE improves TextVQA by 1.60 points and GQA attribute accuracy by 1.31 points, achieving a 3.3× speedup over SmolVLM2-2.2B with only 16 extra local tokens.
By Amitangshu Mukherjee, Kaushik Roy
arXiv:2606.18974v3 Announce Type: replace
Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly a...
By Pengyu Li, Zhitao Gao, Lingling Zhang, Muye Huang, Yuanming Li, Fangzhi Xu, Jun Liu
arXiv:2510. 00054v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding tasks.
By Xianjie Liu, Yiman Hu, Yixiong Zou, Liang Wu, Jian Xu, Bo Zheng
LatentPress compresses conversational histories and long documents into continuous memory tokens that a frozen decoder can read directly, eliminating the need for text reconstruction at inference. The method achieves 4–16× compression with only a small adapter (0.1% of the decoder’s parameters) and outperforms text summaries and OCR-based compression on LongMemEval and LongBench-QA benchmarks. Writing and reading are significantly faster than traditional text summarization or OCR reconstruction, demonstrating a practical machine-facing context interface beyond text and vision.
By Zhengze Zhou, Hejian Sang
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency.
The paper introduces VIG (Visual Information Gain), an information‑theoretic reward that evaluates each token in a multimodal chain‑of‑thought by measuring how much the image reduces its predictive uncertainty. VIG is computed online using two forward passes—one with and one without the image—eliminating the need for reference chains or external annotations. Experiments on six multimodal reasoning benchmarks and multiple Qwen3‑VL‑Thinking model sizes show that VIG consistently improves the accuracy–efficiency trade‑off, demonstrating that efficient multimodal reasoning arises from increasing visual information density rather than merely limiting chain length.
By Wen Luo, Xiaohan Yi, Xiaotao Huang, Liqun Huang
arXiv:2609.16722v1 Announce Type: new
Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates co...
By Haoyu Guo, Yuan Feng, Junlin Lv, Mingjun Xiao, S Kevin Zhou, Xike Xie
Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference cost from multi-step diffusion.
The paper introduces Visual Retrieval Heads (VRHs), a small fraction of attention heads in vision‑language models that are causally responsible for grounding text descriptions to image regions. By recasting head‑scoring methods and evaluating across eleven VLMs and five benchmarks, the authors show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect. VRHs generalize across various visual reference tasks, preserve output format while corrupting localization, and transfer causally across models sharing an LLM backbone.
By Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung
arXiv:2608. 04726v1 Announce Type: new Abstract: Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels.
By Yongxin Wang, Ruizhe Zhou, Yueling Tang, Yingying Zhu, Xuemin Zhao, Xiaojun Chang, Xiaodan Liang