arXiv:2606. 14765v1 Announce Type: cross Abstract: Self-supervised video representation learning has recently advanced through contrastive learning, masked reconstruction, and predictive representation learning.
By Qinwu Xu
The study investigates how the geometry of representations in artificial neural networks can be steered to improve bidirectional alignment with biological neural responses. By applying spectral regularization during training of self‑supervised contrastive vision models, the authors increased reverse predictivity by 55% while only modestly reducing forward predictivity. The changes also lowered effective dimensionality and reorganized the shared subspace, making forward and reverse predictivity more symmetric at certain spectral exponents.
By Samuel Kostousov, Abhinn Kaushik, Brokoslaw Laschowski
The study investigates how aligning the representational geometry of artificial neural networks can improve bidirectional predictivity with biological neural responses. By applying spectral regularization during training of self‑supervised contrastive vision models, the authors increased reverse predictivity by 55% while only modestly reducing forward predictivity. The adjustments also lowered effective dimensionality and reorganized the shared representational subspace, making forward and reverse predictivity more symmetric at intermediate spectral exponents.
arXiv:2606. 05109v1 Announce Type: new Abstract: To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information.
By Vasiliki Rizou, Pascal Frossard, Dorina Thanou
arXiv:2609.06460v1 Announce Type: new
Abstract: Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various...
By Qi Zhang, Runyu Zhou, Yifei Wang, Yisen Wang
arXiv:2609.23881v1 Announce Type: new
Abstract: Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction....
By Markus Karmann, Shile Li, Christian Intern\`o, Bruno Andreis, David Klindt, Randall Balestriero, Jindong Gu, Philip Torr, Qi Zhang, Peng-Tao Jiang, Hao Zhang, Bo Li, Onay Urfalioglu
arXiv:2609.40347v1 Announce Type: new
Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...
By Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information. Learning disentangled representations is a principled way to identify these underlying shared and unique factors that are hidden in observational data.
arXiv:2502. 14424v3 Announce Type: replace-cross Abstract: Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified.
By Yuling Jiao, Wensen Ma, Defeng Sun, Hansheng Wang, Yang Wang
arXiv:2603. 15263v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) has revolutionized representation learning, with Joint-Embedding Architectures (JEAs) emerging as an effective approach for capturing semantic features.
By Konstantinos Almpanakis, Anna Kreshuk
arXiv:2610.01589v1 Announce Type: new
Abstract: Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied syst...
By Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni
arXiv:2608.24093v1 Announce Type: cross
Abstract: Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining...
By Jheng-Ling Lee, Shang-Tse Chen