arXiv Machine Learning

Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures

The paper investigates the use of the squared norm of a whitened foundation‑model embedding as a training‑free likelihood surrogate. It shows that the apparent Gaussianity of whitened coordinates stems from the projection central limit theorem, not from a true joint Gaussian distribution, and that the norm is systematically over‑dispersed compared to a Gaussian reference. The authors explain that whitening reverses the encoder’s spectral hierarchy, concentrating norm contributions in near‑degenerate directions dominated by noise, and propose interpreting the squared norm as a Mahalanobis measure of semantic atypicality rather than a log‑likelihood.

arXiv Machine Learning
Jun 16

InfoNCE Induces Gaussian Distribution

arXiv:2602. 24012v2 Announce Type: replace Abstract: Contrastive learning has become a cornerstone of modern representation learning, allowing training with massive unlabeled data for both task-specific and general (foundation) models.

By Roy Betser, Eyal Gofer, Meir Yossef Levi, Guy Gilboa
arXiv Machine Learning
Sep 18

SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption

SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption proposes a new method for handling sample-heterogeneous corruption in multimodal training. The authors prove that batch-level, sample-agnostic linear estimators with a shared modulation parameter inevitably incur bias, and that a sample-level all-or-nothing gating strategy is the only unbiased approach within a natural estimator class. SAGG implements a binary retain-or-discard decision per sample using an online feature-norm quality test and a truncation mechanism for variance control, and demonstrates convergence to clean-loss stationary points while achieving superior performance over ten existing methods on Kinetics-Sounds and UCF-101 under various corruption scenarios.

By Wentao Zhang, Yifan Zhu, Yutong Zhang, Wentao Mo
arXiv Machine Learning
Jun 24

The Degeneracy Distillery

arXiv:2606. 23838v1 Announce Type: new Abstract: When two or more parameters or labels produce similar data, they are degenerate, or hard to distinguish.

By T. Lucas Makinen, Deaglan J. Bartlett, Niall Jeffrey, Benjamin D. Wandelt
arXiv Machine Learning
Sep 21

On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation

The paper investigates the limits of the maximal coding rate reduction (MCR²) framework for out‑of‑distribution (OOD) generalisation. It shows that MCR² can lead to complete prediction failure under distribution shift, even when a perfectly stable feature is available, and that adding invariance principles from IRM or REx does not resolve this issue. The authors conclude that additional assumptions or learning principles are needed to guarantee stable OOD predictions with MCR².

By Menghui Zhou, Gaoshan Bi, Vitaveska Lanfranchi, Po Yang