Conditional Predictive Sufficient Statistics for Visual Representation Learning
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces Conditional Predictive Sufficient Statistics (CPSS) as a formal way to capture useful visual representations that preserve latent factors shared with future data while discarding noise. It shows that predicting the next image patch embedding with a cosine loss approximates maximum likelihood under a von Mises-Fisher model, and that stop‑gradient alone does not enforce sufficiency. Experiments on MNIST and CIFAR‑10 with small causal Transformers demonstrate that CPSS readouts outperform intermediate blocks, while removing stop‑gradient collapses embedding rank even when the pretext loss appears perfect.
arXiv:2609.38485v1 Announce Type: new Abstract: Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known t...
arXiv:2608. 25138v1 Announce Type: new Abstract: Stochastic masking, cropping, or modality removal makes deterministic reconstruction an incomplete target: one observation can admit many clean completions.
arXiv:2608. 04879v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is independently parameterized.
arXiv:2604.05819v2 Announce Type: replace-cross Abstract: Interpreting the decisions of complex computer vision models is crucial to establish trust and accountability, especially in safety-critical...
arXiv:2608. 12939v1 Announce Type: new Abstract: Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance.