arXiv Computer Vision
6d ago

Conditional Predictive Sufficient Statistics for Visual Representation Learning

The paper introduces Conditional Predictive Sufficient Statistics (CPSS) as a formal way to capture useful visual representations that preserve latent factors shared with future data while discarding noise. It shows that predicting the next image patch embedding with a cosine loss approximates maximum likelihood under a von Mises-Fisher model, and that stop‑gradient alone does not enforce sufficiency. Experiments on MNIST and CIFAR‑10 with small causal Transformers demonstrate that CPSS readouts outperform intermediate blocks, while removing stop‑gradient collapses embedding rank even when the pretext loss appears perfect.

By Yuzhou Hong