arXiv Computer Vision

EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

EviViT is a lightweight attachment for pretrained vision transformers that learns where to focus detail in high‑resolution images. It uses human visual‑search traces to supervise a question‑conditioned evidence density, guiding regional re‑reading and efficient visual token allocation. The method connects regional features to the global scene via a sparse, coordinate‑aware bridge, improving fine‑grained accuracy across nine host models while using fewer tokens than global‑only processing.

arXiv Computer Vision
Sep 1

State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models

arXiv:2608.28698v1 Announce Type: new Abstract: Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model...

By Mingxu Chai, Chenyu Liu, Ziyu Shen, Jiazheng Zhang, Kaidi Zhang, Ruoyu Chen, Jun Long, Jihua Kang, Tao Gui, Qi Zhang
arXiv Computer Vision
Sep 18

Region-Level Policy Optimization for Fine-grained MLLM Perception

The paper introduces Vision‑RL2, a region‑level reinforcement learning approach that optimizes a lightweight proposal network for fine‑grained multimodal large language model (MLLM) perception. By treating coherent image regions as actions and scoring them with a frozen MLLM reader, the method selectively focuses visual resolution on evidence, reducing token usage while improving accuracy across multiple benchmarks and backbones. The approach eliminates the need for region annotations, response sampling, or reasoning trajectories, and the refined proposals enable sparse encoding that magnifies relevant evidence.

By Yuheng Shi, Xiaohuan Pei, Minjing Dong, Chang Xu
arXiv Computer Vision
Aug 28

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

The paper introduces Visual Retrieval Heads (VRHs), a small fraction of attention heads in vision‑language models that are causally responsible for grounding text descriptions to image regions. By recasting head‑scoring methods and evaluating across eleven VLMs and five benchmarks, the authors show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect. VRHs generalize across various visual reference tasks, preserve output format while corrupting localization, and transfer causally across models sharing an LLM backbone.

By Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung