A Theoretical Framework for Masked Pretraining (MPT)
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 05173v1 Announce Type: cross Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are strongly anchored to surface-form token identity rather than deeper semantic structure.
arXiv:2608. 08309v1 Announce Type: cross Abstract: We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy.
arXiv:2603. 15263v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) has revolutionized representation learning, with Joint-Embedding Architectures (JEAs) emerging as an effective approach for capturing semantic features.
Self-supervision is a powerful technique for learning visual representations from unlabeled data. Existing techniques primarily adopt a two-stage approach for self-supervised learning (SSL): a pretraining stage on unlabeled data followed by a finetuning stage on labeled data.
The study investigates how the geometry of representations in artificial neural networks can be steered to improve bidirectional alignment with biological neural responses. By applying spectral regularization during training of self‑supervised contrastive vision models, the authors increased reverse predictivity by 55% while only modestly reducing forward predictivity. The changes also lowered effective dimensionality and reorganized the shared subspace, making forward and reverse predictivity more symmetric at certain spectral exponents.
The study investigates how aligning the representational geometry of artificial neural networks can improve bidirectional predictivity with biological neural responses. By applying spectral regularization during training of self‑supervised contrastive vision models, the authors increased reverse predictivity by 55% while only modestly reducing forward predictivity. The adjustments also lowered effective dimensionality and reorganized the shared representational subspace, making forward and reverse predictivity more symmetric at intermediate spectral exponents.