arXiv Computer Vision

On the Role of the Projector in Contrastive Self-Supervised Learning: Last-Layer Rank Dynamics Drive Representation Quality

The paper investigates the function of the projector in contrastive self‑supervised learning by analyzing rank dynamics in the encoder and projector. It finds that rank reduction mainly occurs in the last layer and proposes a targeted weight regularization for that layer, which outperforms global orthogonal regularization. Experiments on SimCLR with ImageNet100 and CIFAR datasets show more than 1% Top‑1 accuracy improvement and consistent gains over baseline variants.

arXiv Machine Learning
1d ago

$\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

The paper introduces SACReg, a spectral anti-collapse regularizer that enforces λ-balance across weight matrices to prevent dimensional collapse in the backbone of joint-embedding self-supervised learning models. Applied to JEPA, the resulting λ-JEPA improves ImageNet-1k classification and linear-probe transfer on eight image datasets, and also outperforms prior video self-supervised methods on Something-Something-v2 and Kinetics-400.

By Berker Demirel, Cl\'ementine Domin\'e, Valentino Maiorca, Marco Fumero, Marco Mondelli, Francesco Locatello
Hugging Face Trending Papers
Jul 2

DRDN: Decoupled Representation Dynamic Network for From-Scratch ViT Class-Incremental Learning

Dynamic expansion methods for class-incremental learning (CIL) protect task-specific knowledge by growing dedicated tokens or subnetworks, yet our analyses suggest that classification supervision alone does not sufficiently preserve task-agnostic shared backbone representations over long incremental sequences. We identify two intertwined challenges: cross-task confusion from sequential training on predominantly current-task data, which biases decision boundaries toward recent tasks; and under-optimized shared representations in the backbone that cap long-term discriminability as tasks accumulate.

arXiv Computer Vision
Sep 3

Breaking the Geometric Bottleneck: Contrastive Expansion in Asymmetric Cross-Modal Distillation

The paper investigates how knowledge distillation from Vision Transformers to smaller CNNs can cause dimensional collapse in the student’s representation space. Using SVD and Shannon entropy, the authors show that cosine‑based distillation leads to a drastic reduction in effective rank, while adding an InfoNCE objective can double the rank but harms downstream accuracy due to signal dilution. They further demonstrate that a label‑aware contrastive objective (Supervised Contrastive distillation) can maintain or improve accuracy without unnecessary rank expansion, indicating that effective rank alone is not a reliable indicator of representation quality.

By Kabir Thayani
Hugging Face Trending Papers
Jul 15

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.