Self-Improving Small Object Grounding in LVLMs
arXiv:2606. 01612v1 Announce Type: cross Abstract: Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning?
Robust visual classification often depends on localizing the main foreground objects in an image while ignoring contextual distractors. Surprisingly, we find that the attention maps of smaller self-supervised ViTs localize foreground objects better than those of larger ViTs.
arXiv:2606. 01612v1 Announce Type: cross Abstract: Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning?
CoViT introduces a self‑supervised framework that enhances Vision Transformers with instance‑aware representations by leveraging geometry‑guided contrastive learning. It refines attention maps to generate instance masks and constructs triplets that mine the hardest intra‑ and inter‑instance examples, driving a contrastive loss that reduces intra‑instance variance while increasing inter‑instance margins. The method yields consistent AP gains of over 2 points on instance‑level tasks without requiring extra decoders or labels.
arXiv:2605.12491v2 Announce Type: replace Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design im...
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in light...
arXiv:2609.37659v1 Announce Type: cross Abstract: There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit....
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction c...
arXiv:2608.22996v1 Announce Type: new Abstract: Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution s...
arXiv:2505.16157v3 Announce Type: replace Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Tran...
arXiv:2607. 13395v1 Announce Type: new Abstract: The pursuit of autonomously self-improving models has attracted growing interest in the era of large-scale foundation models.
arXiv:2607.09086v2 Announce Type: replace Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
arXiv:2507.00754v3 Announce Type: replace Abstract: The integration of Large Language Model (LLMs) blocks with Vision Transformers (ViTs) holds immense promise for vision-only tasks by leveraging the...
arXiv:2407.03463v2 Announce Type: replace-cross Abstract: In the realm of self-supervised learning (SSL), conventional wisdom has gravitated towards the utility of massive, general domain datasets fo...