arXiv Machine Learning By Hugues Van Assel, Edward De Brouwer, Saeed Saremi, Gabriele Scalia, Aviv Regev

Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation

Read the original on arXiv Machine Learning →

arXiv:2606. 00514v1 Announce Type: new Abstract: Generative modeling and self-supervised representation learning (SSL) optimize structurally different objectives: generative training rewards distributional fidelity, while SSL rewards semantic coherence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 7

An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders

The paper investigates whether pretrained image models can generalize to unseen datasets by clustering their embeddings. Using encoders trained only on ImageNet‑1k, both supervised and self‑supervised, the authors evaluate clustering performance on out‑of‑domain images. They find that supervised encoders perform better within the training domain, while self‑supervised encoders excel far outside it, and that fine‑tuning self‑supervised models reverses this trend. Additionally, the study shows that the silhouette score in UMAP‑reduced space correlates strongly with clustering accuracy, offering a proxy metric when labels are unavailable.

By Scott C. Lowe, Joakim Bruslund Haurum, Sageev Oore, Thomas B. Moeslund, Graham W. Taylor
arXiv Computer Vision
Aug 27

DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors

DEFUSE is a backdoor detection framework for self‑supervised encoders that uses a conditional diffusion generative model to estimate representation‑conditioned image likelihoods. By fine‑tuning a pretrained diffusion model, DEFUSE performs semantic reconstruction in a reference encoder’s representation space, enabling it to detect backdoors without needing uninfected data or precomputed pseudo‑labels. Experiments show that DEFUSE outperforms existing detectors on both visual SSL and vision‑language encoders, reducing reliance on prior knowledge of the victim model or attack strategy.

By Tuo Chen, Jie Gui, Minjing Dong, Lanting Fang, Ju Jia, Benlei Cui, Jian Liu
arXiv AI
Sep 21

The Impact of Semantic Pairs on Self-Supervised Representation Learning

The paper investigates the effect of using semantic positive pairs—different instances of the same class—in self‑supervised visual representation learning. By creating matched ImageNet‑1K subsets of augmented pairs and manually curated semantic pairs, the authors compare contrastive and non‑contrastive SSL methods under identical training conditions. Across transfer learning and object detection tasks, semantic‑pair pretraining consistently outperforms augmented‑pair pretraining, with contrastive methods like SimCLR showing the largest gains, indicating that semantic pairs foster additional invariances beyond standard augmentations.

By Mohammad Alkhalefi, Georgios Leontidis, Mingjun Zhong
arXiv AI
Sep 24

AdaDim: Dimensionality Adaptation for SSL Representational Dynamics

AdaDim introduces a training strategy for self‑supervised learning that adaptively balances dimensionality increase and mutual information reduction. By gradually regularizing the projection head while encouraging feature decorrelation and sample uniformity, AdaDim achieves up to 3% performance gains over standard SSL baselines without relying on costly techniques such as queues or predictor networks. The method demonstrates that optimal SSL models do not simply maximize dimensionality or minimize mutual information, but find a trade‑off between the two.

By Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib