Hugging Face Trending Papers

Scalable Patch-Level Self-Supervised Learning

arXiv Computer Vision
1d ago

Scalable Patch-Level Self-Supervised Learning

arXiv:2610.10013v1 Announce Type: new Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of mul...

By Maximilian Seitzer, Gabriele Trivigno, Anton\'in Vobeck\'y, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Sim\'eoni, Piotr Bojanowski
arXiv AI
Oct 1

Which Tasks Survive Self-Supervised Learning?

arXiv:2609.38393v1 Announce Type: cross Abstract: Same-instance self-supervised learning (SSL) learns representations by enforcing consistency across two views of the same underlying instance. This p...

By Achleshwar Luthra, Lucas Bryant, Tracy Zhu, Tomer Galanti
arXiv Machine Learning
Jul 7

Self-Supervised Learning from Structural Invariance

arXiv:2602. 02381v2 Announce Type: replace Abstract: Joint-embedding self-supervised learning (SSL), the key paradigm for unsupervised representation learning from visual data, learns from invariances between semantically-related data pairs.

By Yipeng Zhang, Hafez Ghaemi, Jungyoon Lee, Shahab Bakhtiari, Eilif B. Muller, Laurent Charlin
arXiv AI
Sep 24

AdaDim: Dimensionality Adaptation for SSL Representational Dynamics

AdaDim introduces a training strategy for self‑supervised learning that adaptively balances dimensionality increase and mutual information reduction. By gradually regularizing the projection head while encouraging feature decorrelation and sample uniformity, AdaDim achieves up to 3% performance gains over standard SSL baselines without relying on costly techniques such as queues or predictor networks. The method demonstrates that optimal SSL models do not simply maximize dimensionality or minimize mutual information, but find a trade‑off between the two.

By Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib
arXiv Computer Vision
Sep 28

A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources

The paper reports a controlled study of self‑supervised learning (SSL) objectives for image and video pretraining under limited data, architecture, and compute budgets. It compares contrastive, reconstruction, feature‑prediction, and diffusion methods, finding that DINOv2‑style pretraining delivers the best overall performance. Combining DINOv2 with video SSL objectives such as VideoMAE improves image classification and segmentation but harms video tracking and camera‑pose estimation, highlighting a trade‑off between semantic and geometric learning.

By Brun\'o B. Englert, Gijs Dubbelman