arXiv Machine Learning

Divergence-Based Similarity Function for Multi-View Contrastive Learning

The paper introduces a divergence-based similarity function (DSF) for multi-view contrastive learning, representing each set of augmented views as a distribution and measuring similarity via distribution divergence. DSF captures joint structure across all views, outperforming prior pairwise methods on tasks such as kNN classification, linear evaluation, transfer learning, and distribution shift. It also offers greater efficiency and eliminates the need for a temperature hyperparameter, unlike cosine similarity.

arXiv Computer Vision
Sep 11

Prototype Matters: Modality-unified Prototype Self-distillation for Unsupervised Visible-infrared Person Re-identification

The paper introduces a new framework for unsupervised visible‑infrared person re‑identification that leverages modality‑unified prototypes. By contrasting with prototypes that unify both modalities, the method jointly optimizes similarity within and across modalities, improving modality invariance. A self‑distillation step refines instance‑prototype relationships using a steady teacher, resulting in a simple yet effective model validated on standard VI‑ReID benchmarks.

By Menglin Wang, Xiaojin Gong
arXiv AI
Aug 20

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.

By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
arXiv Machine Learning
Jun 16

InfoNCE Induces Gaussian Distribution

arXiv:2602. 24012v2 Announce Type: replace Abstract: Contrastive learning has become a cornerstone of modern representation learning, allowing training with massive unlabeled data for both task-specific and general (foundation) models.

By Roy Betser, Eyal Gofer, Meir Yossef Levi, Guy Gilboa
arXiv AI
Sep 2

ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives

ViTAMINS is a method that incorporates synthetic hard negatives into unsupervised vision transformer pretraining to enhance representation quality. The approach is evaluated on ImageNet and a range of downstream tasks—including transfer learning, image retrieval, copy detection, and image/video segmentation—showing significant performance gains. The synthetic negatives also lead to emergent properties, such as representations that encode explicit semantic information and act as strong classifiers, improving over baselines by up to 11.3%.

By Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos, Tania Stathaki