arXiv Machine Learning

Understanding Self-Supervised Learning via Latent Distribution Matching

arXiv:2605. 03517v4 Announce Type: replace Abstract: Self-supervised learning (SSL) excels at finding general-purpose latent representations from complex data, yet lacks a unifying theoretical framework that explains the diverse existing methods and guides the design of new ones.

arXiv Machine Learning
Jul 7

Self-Supervised Learning from Structural Invariance

arXiv:2602. 02381v2 Announce Type: replace Abstract: Joint-embedding self-supervised learning (SSL), the key paradigm for unsupervised representation learning from visual data, learns from invariances between semantically-related data pairs.

By Yipeng Zhang, Hafez Ghaemi, Jungyoon Lee, Shahab Bakhtiari, Eilif B. Muller, Laurent Charlin
arXiv Machine Learning
Aug 11

A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization

arXiv:2602. 10680v2 Announce Type: replace-cross Abstract: Many real-world datasets contain hidden structure that cannot be detected by simple linear correlations between input features.

By Vicente Conde Mendes, Lorenzo Bardone, C\'edric Koller, Jorge Medina Moreira, Vittorio Erba, Emanuele Troiani, Lenka Zdeborov\'a
Hugging Face Trending Papers
Jun 29

Simplifying Flow Matching Transformations with Low-Rank Mixture Models

Normalizing flows are powerful generative models that learn an invertible mapping between complex data distributions and simple latent distributions, typically a standard normal density. However, this choice of latent density can impose unnecessary complexity on the learned flow transformation due to the topological mismatch between the latent and data densities, leading to slower training and suboptimal performance.

arXiv Machine Learning
Jul 27

Unbiased Open World Regularization for Fair Self-Supervised Learning

arXiv:2607. 22149v1 Announce Type: new Abstract: Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spurious biases in the dataset.

By L{\'e}o Nicollier (CB, ATT), Marc Pic (ATT), Pablo Mus{\'e} (CB, IFUMI), Enric Meinhardt-Llopis (CB), Gabriele Facciolo (CB)