Hugging Face Trending Papers

$A^2$: Smaller Self-Supervised ViTs Localize Better than Larger Ones

Read the original on Hugging Face Trending Papers →

Robust visual classification often depends on localizing the main foreground objects in an image while ignoring contextual distractors. Surprisingly, we find that the attention maps of smaller self-supervised ViTs localize foreground objects better than those of larger ViTs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Sep 3

CoViT: Instance-Correspondence Contrastive Learning for Vision Transformer

CoViT introduces a self‑supervised framework that enhances Vision Transformers with instance‑aware representations by leveraging geometry‑guided contrastive learning. It refines attention maps to generate instance masks and constructs triplets that mine the hardest intra‑ and inter‑instance examples, driving a contrastive loss that reduces intra‑instance variance while increasing inter‑instance margins. The method yields consistent AP gains of over 2 points on instance‑level tasks without requiring extra decoders or labels.

By Yisen Wang, Zhirong Wu, Limin Wang
arXiv Machine Learning
4d ago

Are In-Context Images Worth 10 Dimensions?

arXiv:2609.37659v1 Announce Type: cross Abstract: There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit....

By Adhemar de Senneville, Xavier Bou, J\'er\'emy Anger, Rafael Grompone, Gabriele Facciolo