arXiv Machine Learning

$\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

The paper introduces SACReg, a spectral anti-collapse regularizer that enforces λ-balance across weight matrices to prevent dimensional collapse in the backbone of joint-embedding self-supervised learning models. Applied to JEPA, the resulting λ-JEPA improves ImageNet-1k classification and linear-probe transfer on eight image datasets, and also outperforms prior video self-supervised methods on Something-Something-v2 and Kinetics-400.

arXiv Computer Vision
Sep 23

On the Role of the Projector in Contrastive Self-Supervised Learning: Last-Layer Rank Dynamics Drive Representation Quality

The paper investigates the function of the projector in contrastive self‑supervised learning by analyzing rank dynamics in the encoder and projector. It finds that rank reduction mainly occurs in the last layer and proposes a targeted weight regularization for that layer, which outperforms global orthogonal regularization. Experiments on SimCLR with ImageNet100 and CIFAR datasets show more than 1% Top‑1 accuracy improvement and consistent gains over baseline variants.

By Siladittya Manna, Priyangshu Mandal, Umapada Pal, Saumik Bhattacharya
arXiv AI
Aug 28

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.

By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
arXiv Machine Learning
Sep 7

An Analysis of Self-supervised Pre-training with Dependent Samples

The paper investigates self‑supervised pre‑training that uses multiple data augmentations of the same unlabeled sample. It shows that pooling these dependent augmentations together yields statistical estimation error bounds that are never worse than, and sometimes better than, partitioning the data into independent subsets. The analysis explains why using many augmentations is practically advantageous, especially when their correlations have mild effects or reduce estimation variance.

By Maximilian Fleissner, Debarghya Ghoshdastidar, Samory Kpotufe