arXiv Machine Learning By Kawtar Zaher, Ilyass Moummad, Olivier Buisson, Alexis Joly

Self-Supervised Learning as Discrete Communication

Read the original on arXiv Machine Learning →

arXiv:2602. 09764v2 Announce Type: replace-cross Abstract: Most self-supervised learning (SSL) methods learn continuous visual representations by aligning different views of the same input, offering limited control over how information is structured across representation dimensions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 7

Self-Supervised Learning from Structural Invariance

arXiv:2602. 02381v2 Announce Type: replace Abstract: Joint-embedding self-supervised learning (SSL), the key paradigm for unsupervised representation learning from visual data, learns from invariances between semantically-related data pairs.

By Yipeng Zhang, Hafez Ghaemi, Jungyoon Lee, Shahab Bakhtiari, Eilif B. Muller, Laurent Charlin
arXiv Computer Vision
Sep 24

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.

By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv AI
Sep 21

The Impact of Semantic Pairs on Self-Supervised Representation Learning

The paper investigates the effect of using semantic positive pairs—different instances of the same class—in self‑supervised visual representation learning. By creating matched ImageNet‑1K subsets of augmented pairs and manually curated semantic pairs, the authors compare contrastive and non‑contrastive SSL methods under identical training conditions. Across transfer learning and object detection tasks, semantic‑pair pretraining consistently outperforms augmented‑pair pretraining, with contrastive methods like SimCLR showing the largest gains, indicating that semantic pairs foster additional invariances beyond standard augmentations.

By Mohammad Alkhalefi, Georgios Leontidis, Mingjun Zhong
Hugging Face Trending Papers
Jun 3

RePercENT: Scaling Disentangled Representation Learning Beyond Two Modalities

To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information. Learning disentangled representations is a principled way to identify these underlying shared and unique factors that are hidden in observational data.