arXiv Machine Learning By Berker Demirel, Cl\'ementine Domin\'e, Valentino Maiorca, Marco Fumero, Marco Mondelli, Francesco Locatello

$\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

Read the original on arXiv Machine Learning →

The paper introduces SACReg, a spectral anti-collapse regularizer that enforces λ-balance across weight matrices to prevent dimensional collapse in the backbone of joint-embedding self-supervised learning models. Applied to JEPA, the resulting λ-JEPA improves ImageNet-1k classification and linear-probe transfer on eight image datasets, and also outperforms prior video self-supervised methods on Something-Something-v2 and Kinetics-400.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 23

On the Role of the Projector in Contrastive Self-Supervised Learning: Last-Layer Rank Dynamics Drive Representation Quality

The paper investigates the function of the projector in contrastive self‑supervised learning by analyzing rank dynamics in the encoder and projector. It finds that rank reduction mainly occurs in the last layer and proposes a targeted weight regularization for that layer, which outperforms global orthogonal regularization. Experiments on SimCLR with ImageNet100 and CIFAR datasets show more than 1% Top‑1 accuracy improvement and consistent gains over baseline variants.

By Siladittya Manna, Priyangshu Mandal, Umapada Pal, Saumik Bhattacharya
arXiv AI
Aug 28

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.

By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner