Conditional Predictive Sufficient Statistics for Visual Representation Learning
Read the original on arXiv Computer Vision →The paper introduces Conditional Predictive Sufficient Statistics (CPSS) as a formal way to capture useful visual representations that preserve latent factors shared with future data while discarding noise. It shows that predicting the next image patch embedding with a cosine loss approximates maximum likelihood under a von Mises-Fisher model, and that stop‑gradient alone does not enforce sufficiency. Experiments on MNIST and CIFAR‑10 with small causal Transformers demonstrate that CPSS readouts outperform intermediate blocks, while removing stop‑gradient collapses embedding rank even when the pretext loss appears perfect.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.