arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu
arXiv:2608.28216v1 Announce Type: new
Abstract: Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is ab...
By Kishor Datta Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Md. Sadman Haque, Mohd Ariful Haque
arXiv:2609.13225v1 Announce Type: cross
Abstract: Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fail...
By Sarthak Sattigeri
arXiv:2607.05568v2 Announce Type: replace-cross
Abstract: Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods...
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
arXiv:2608. 07088v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive.
By Qiyanhui Lu, Han Wu, Rongjian Xu, Tingzhang Luo, Cheng Fan, Xinghao Chen, Minjing Dong, Jufeng Yang, Jianyuan Guo
arXiv:2606. 14757v1 Announce Type: cross Abstract: Though Vision Transformers (ViTs) have become the dominant backbone in many computer vision tasks, due to permutation equivariance, their attention mechanism lacks explicit spatial inductive biases.
By Leyla Naz Candogan, Arshia Afzal, Pol Puigdemont, Volkan Cevher
The paper investigates why vision‑language models like LLaVA‑1.5‑7B hallucinate objects in captions and proposes a targeted fix. By ranking attention heads whose image attention drops around hallucinated words, the authors identify 32 key heads and apply a head‑sliced LoRA adapter plus an inference‑time grounding controller. On COCO images, this combined method reduces hallucinated captions from 37% to 23% and hallucinated object mentions from 15.6% to 9.6%, while also lowering object recall.
By Armaan Sandhu, Abhilasha Senapati, Hima Kammachi
arXiv:2608. 04190v1 Announce Type: new Abstract: Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not recover it: combiners such as majority voting trade recall for precision and are brittle to coordinated failures.
By Mario Leiva, Yue Ma, Qinru Qiu, Gerardo Simari, Paulo Shakarian
arXiv:2608.22066v1 Announce Type: cross
Abstract: Attention-based multiple instance learning (ABMIL) using pathology foundation model embeddings is effective for slide-level tasks, but exhaustive inf...
By Duncan Stothers, Ren-Chin Wu, William Lotter
arXiv:2608.23253v1 Announce Type: cross
Abstract: Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Exi...
By Taoyu Qian, Qi Wang, Daqian Shi, Yuanhao Jiang, Shang Gao, Hualong Yu
arXiv:2609.06232v1 Announce Type: cross
Abstract: Ground-truth defect masks in industrial inspection datasets are typically reserved for evaluation. This paper repurposes them as spatial supervision...
By Sajjad Rezvani Boroujeni, Muskan Saraf, Gnana Tulasi Makineni, Tom Bush, Hossein Abedi
ReVisIT is a train‑free framework that turns retrieved image‑label pairs into units of visual thought, combining structured class definitions, multimodal retrieval, and alternating user/assistant injection before joint decoding. On several benchmarks—including Fast Open MiniImageNet, Bongard‑OpenWorld, and the newly released MAAC‑Bench—ReVisIT achieves performance comparable to or surpassing large, trained models while using far fewer parameters. The approach demonstrates that high‑quality retrieval and a simple turns layer can provide a universal performance boost across diverse multimodal tasks.
By Bingchen Huang, Zhiling Wang, Yifu Chen, Yuanchao Du