Hugging Face Trending Papers

Conditional Predictive Sufficient Statistics for Visual Representation Learning

arXiv Computer Vision
6d ago

Conditional Predictive Sufficient Statistics for Visual Representation Learning

The paper introduces Conditional Predictive Sufficient Statistics (CPSS) as a formal way to capture useful visual representations that preserve latent factors shared with future data while discarding noise. It shows that predicting the next image patch embedding with a cosine loss approximates maximum likelihood under a von Mises-Fisher model, and that stop‑gradient alone does not enforce sufficiency. Experiments on MNIST and CIFAR‑10 with small causal Transformers demonstrate that CPSS readouts outperform intermediate blocks, while removing stop‑gradient collapses embedding rank even when the pretext loss appears perfect.

By Yuzhou Hong
arXiv Machine Learning
Sep 22

Common Cause, Not Cross-Attention: Blocking Visual Shortcuts in Audio-Video Generation

The paper investigates how joint audio–video generation models learn to associate sound with visual appearance rather than the underlying causal event, a problem termed the visual shortcut. By constructing a structural causal model where audio is independent of video appearance, the authors demonstrate that cross‑attention and shared latent approaches fail when appearance‑event correlations are broken, and that common‑cause routing does not solve the issue. They propose blocking the shortcut via interventions on nuisance variables, proving that counterfactual invariance is necessary and sufficient for identifying the causal predictor, and validate this approach on synthetic and real datasets, including a pretrained video‑to‑audio generator.

By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv AI
Jul 21

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation

arXiv:2607. 18042v1 Announce Type: cross Abstract: End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly train the policy state to be predictive of future visual outcomes.

By Lingfeng Zhang, Zhanguang Zhang, Liheng Ma, Tongtong Cao, Yingxue Zhang