The dimensional collapse of representations in self-supervised contrastive learning is an ever-present issue. One notable technique to prevent such a collapse of representations is using a multi-layer...
The paper investigates the function of the projector in contrastive self‑supervised learning by analyzing rank dynamics in the encoder and projector. It finds that rank reduction mainly occurs in the last layer and proposes a targeted weight regularization for that layer, which outperforms global orthogonal regularization. Experiments on SimCLR with ImageNet100 and CIFAR datasets show more than 1% Top‑1 accuracy improvement and consistent gains over baseline variants.
By Siladittya Manna, Priyangshu Mandal, Umapada Pal, Saumik Bhattacharya
arXiv:2607. 07513v1 Announce Type: new Abstract: Self-supervised learning matches supervised accuracy from a fraction of the labels, but the labeled-sample efficiency behind this has lacked a theoretical explanation.
By Adam M. Oberman
LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.
By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
arXiv:2608. 15901v1 Announce Type: cross Abstract: Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher values.
By Brian B. Moser, Ahmed Anwar, Tobias Christian Nauen, Shishir Muralidhara, Federico Raue, Ren\'e Schuster, Stanislav Frolov, Andreas Dengel
arXiv:2609.15825v1 Announce Type: new
Abstract: A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's que...
By Dier Tang, Jing Yee Tan, Guangyue Han
The paper investigates self‑supervised pre‑training that uses multiple data augmentations of the same unlabeled sample. It shows that pooling these dependent augmentations together yields statistical estimation error bounds that are never worse than, and sometimes better than, partitioning the data into independent subsets. The analysis explains why using many augmentations is practically advantageous, especially when their correlations have mild effects or reduce estimation variance.
By Maximilian Fleissner, Debarghya Ghoshdastidar, Samory Kpotufe
arXiv:2608. 08309v1 Announce Type: cross Abstract: We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy.
By Nikos Giakoumoglou, Paschalis Giakoumoglou, Tania Stathaki
arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.
By Etienne Boursier, Matthew Bowditch, Matthias Englert, Ranko Lazic
arXiv:2609.06341v1 Announce Type: cross
Abstract: Linear algebra provides the framework of concepts (matrix rank, singular value decomposition (SVD), and eigendecomposition) that modern artificial in...
By Anjaneya Teja Sarma Kalvakolanu
arXiv:2603. 15263v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) has revolutionized representation learning, with Joint-Embedding Architectures (JEAs) emerging as an effective approach for capturing semantic features.
By Konstantinos Almpanakis, Anna Kreshuk
Self-supervision is a powerful technique for learning visual representations from unlabeled data. Existing techniques primarily adopt a two-stage approach for self-supervised learning (SSL): a pretraining stage on unlabeled data followed by a finetuning stage on labeled data.