arXiv Computer Vision

Conditional Predictive Sufficient Statistics for Visual Representation Learning

The paper introduces Conditional Predictive Sufficient Statistics (CPSS) as a formal way to capture useful visual representations that preserve latent factors shared with future data while discarding noise. It shows that predicting the next image patch embedding with a cosine loss approximates maximum likelihood under a von Mises-Fisher model, and that stop‑gradient alone does not enforce sufficiency. Experiments on MNIST and CIFAR‑10 with small causal Transformers demonstrate that CPSS readouts outperform intermediate blocks, while removing stop‑gradient collapses embedding rank even when the pretext loss appears perfect.

arXiv Computer Vision
Aug 27

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.

By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson