arXiv AI

The Impact of Semantic Pairs on Self-Supervised Representation Learning

The paper investigates the effect of using semantic positive pairs—different instances of the same class—in self‑supervised visual representation learning. By creating matched ImageNet‑1K subsets of augmented pairs and manually curated semantic pairs, the authors compare contrastive and non‑contrastive SSL methods under identical training conditions. Across transfer learning and object detection tasks, semantic‑pair pretraining consistently outperforms augmented‑pair pretraining, with contrastive methods like SimCLR showing the largest gains, indicating that semantic pairs foster additional invariances beyond standard augmentations.

arXiv AI
Jun 24

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.

By Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir
arXiv Computer Vision
4d ago

A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources

The paper reports a controlled study of self‑supervised learning (SSL) objectives for image and video pretraining under limited data, architecture, and compute budgets. It compares contrastive, reconstruction, feature‑prediction, and diffusion methods, finding that DINOv2‑style pretraining delivers the best overall performance. Combining DINOv2 with video SSL objectives such as VideoMAE improves image classification and segmentation but harms video tracking and camera‑pose estimation, highlighting a trade‑off between semantic and geometric learning.

By Brun\'o B. Englert, Gijs Dubbelman
arXiv Machine Learning
Jul 7

Self-Supervised Learning from Structural Invariance

arXiv:2602. 02381v2 Announce Type: replace Abstract: Joint-embedding self-supervised learning (SSL), the key paradigm for unsupervised representation learning from visual data, learns from invariances between semantically-related data pairs.

By Yipeng Zhang, Hafez Ghaemi, Jungyoon Lee, Shahab Bakhtiari, Eilif B. Muller, Laurent Charlin
arXiv AI
Jun 16

Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings

arXiv:2606. 15134v1 Announce Type: cross Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched.

By Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja
arXiv Computer Vision
Sep 24

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.

By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun