arXiv Machine Learning

Self-Supervised Representation Learning as Mutual Information Maximization

arXiv Machine Learning
1d ago

Reformulation-Contrastive Learning for Mixed Integer Programs

The paper introduces ReMILP, a reformulation‑contrastive learning framework that uses self‑supervision from equivalent formulations of mixed‑integer linear programs (MILPs). By distinguishing re‑descriptions and substitutions, the method trains a graph neural network and a hypernetwork to predict how variable embeddings transform under changes of variables, achieving invariance and equivariance without solver‑derived labels. The learned representations prove useful for tasks such as binary solution, constraint activity, and integrality gap prediction, and serve as a strong initialization for fine‑tuning.

By Ousema Bouaneni, Mathis Le Bail, Cl\'ement Elliker, Ma\"el Jenny, Sonia Vanier
arXiv Computer Vision
Sep 3

Synergistic Information Disentanglement for Omni-modal Slide Representation Learning in Computational Pathology

The paper introduces Φ-Omni, a self‑supervised learning framework for computational pathology that disentangles synergistic information across histology, genomics, and clinical reports using Partial Information Decomposition. By employing a Synergistic Information Bottleneck and a ΦID objective, the method suppresses redundant signals while maximizing irreducible cross‑modal synergy, leading to improved few‑shot performance on breast and lung whole‑slide image datasets. The authors demonstrate that Φ-Omni outperforms both supervised and other SSL baselines on eight external tasks.

By Mingxin Liu, Chengfei Cai, Anwen Lu, Pengbo Xu, Jun Li, Jinze Li, Depin Chen, Jun Xu
arXiv Machine Learning
Jul 7

Self-Supervised Learning from Structural Invariance

arXiv:2602. 02381v2 Announce Type: replace Abstract: Joint-embedding self-supervised learning (SSL), the key paradigm for unsupervised representation learning from visual data, learns from invariances between semantically-related data pairs.

By Yipeng Zhang, Hafez Ghaemi, Jungyoon Lee, Shahab Bakhtiari, Eilif B. Muller, Laurent Charlin
arXiv Machine Learning
1d ago

Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions

The paper proposes a structured version of the Information Bottleneck (IB) that separates label-relevant structure from within-condition variation using a dual-bottleneck formulation. It introduces a conditional KL term that targets within-condition information, allowing explicit control over nuisance-like variation in learned representations. Experiments demonstrate improved performance in low-data classification and consistent gains on dense prediction tasks.

By Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran