arXiv Machine Learning

An Analysis of Self-supervised Pre-training with Dependent Samples

The paper investigates self‑supervised pre‑training that uses multiple data augmentations of the same unlabeled sample. It shows that pooling these dependent augmentations together yields statistical estimation error bounds that are never worse than, and sometimes better than, partitioning the data into independent subsets. The analysis explains why using many augmentations is practically advantageous, especially when their correlations have mild effects or reduce estimation variance.

arXiv Machine Learning
Jun 10

Post-Training Augmentation Invariance

arXiv:2505. 11702v3 Announce Type: replace Abstract: This work develops a framework for post-training augmentation invariance, in which our goal is to add invariance properties to a pretrained network without altering its behavior on the original, non-augmented input distribution.

By Keenan Eikenberry, Lizuo Liu, Yoonsang Lee
arXiv Machine Learning
1d ago

$\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

The paper introduces SACReg, a spectral anti-collapse regularizer that enforces λ-balance across weight matrices to prevent dimensional collapse in the backbone of joint-embedding self-supervised learning models. Applied to JEPA, the resulting λ-JEPA improves ImageNet-1k classification and linear-probe transfer on eight image datasets, and also outperforms prior video self-supervised methods on Something-Something-v2 and Kinetics-400.

By Berker Demirel, Cl\'ementine Domin\'e, Valentino Maiorca, Marco Fumero, Marco Mondelli, Francesco Locatello
arXiv AI
Jun 15

Learning What to Predict: Downstream-Guided Task Design for Continued Pretraining

arXiv:2601. 22108v2 Announce Type: replace-cross Abstract: Continued pretraining is optimized with fixed self-supervised tasks but selected by downstream performance, creating a coarse feedback loop in which practitioners evaluate checkpoints, change data mixtures or objectives, and restart runs, while individual updates remain blind to target capabilities.

By Shuqi Ke, Giulia Fanti
arXiv AI
Aug 17

Self-Supervised Visual On-Policy Distillation

arXiv:2608. 14144v1 Announce Type: cross Abstract: Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest.

By Yijiang Li, Yijun Liang, Yunjie Tian, Bingyang Wang, Ke Zhang, Zhenfei Yin, Di Fu, Philip Torr, Nuno Vasconcelos
arXiv Computer Vision
Aug 24

When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

arXiv:2608.21098v1 Announce Type: new Abstract: Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or har...

By Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva