arXiv:2606. 01612v1 Announce Type: cross Abstract: Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning?
By Tianze Yang, Yucheng Shi, Ruitong Sun, Ninghao Liu, Jin Sun
CoViT introduces a self‑supervised framework that enhances Vision Transformers with instance‑aware representations by leveraging geometry‑guided contrastive learning. It refines attention maps to generate instance masks and constructs triplets that mine the hardest intra‑ and inter‑instance examples, driving a contrastive loss that reduces intra‑instance variance while increasing inter‑instance margins. The method yields consistent AP gains of over 2 points on instance‑level tasks without requiring extra decoders or labels.
By Yisen Wang, Zhirong Wu, Limin Wang
arXiv:2605.12491v2 Announce Type: replace
Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design im...
By Alan Z. Song, Yinjie Chen, Mu Nan, Deva Ramanan, Michael J. Tarr, Andrew F. Luo
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in light...
arXiv:2609.37659v1 Announce Type: cross
Abstract: There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit....
By Adhemar de Senneville, Xavier Bou, J\'er\'emy Anger, Rafael Grompone, Gabriele Facciolo
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction c...