arXiv Computer Vision By Yaoxin Niu, Zhangquan Chen, Yang Zhang, Xiang An, Zhumei Wang, Chih-Ting Liao, Hongkun Cao, Ruqi Huang

EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

Read the original on arXiv Computer Vision →

EviViT is a lightweight attachment for pretrained vision transformers that learns where to focus detail in high‑resolution images. It uses human visual‑search traces to supervise a question‑conditioned evidence density, guiding regional re‑reading and efficient visual token allocation. The method connects regional features to the global scene via a sparse, coordinate‑aware bridge, improving fine‑grained accuracy across nine host models while using fewer tokens than global‑only processing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.