Hugging Face Trending Papers

Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents

Read the original on Hugging Face Trending Papers →

Latent visual reasoning (LVR) inserts supervised latent tokens between perception and answer generation in vision-language models (VLMs). The field uses alignment between these latents and their visual targets, i.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 10

Reason Through the Latent! Making Latent Visual Reasoning Necessary

The paper introduces Causal Visual Recurrent Reasoning (CVRR), a method that forces multimodal models to rely on latent visual states by using recurrent computation as the sole image‑conditioned path to prediction. CVRR initializes recurrence from the question hidden state after a pretrained vision‑language model has processed the image, repeatedly updates this state while re‑reading the same visual evidence, and removes all other visual traces before decoding. Experiments on several benchmarks show that CVRR maintains strong performance while other latent reasoners lose visual competence, and causal interventions confirm that predictions depend on the recurrent visual trajectory.

By Suhyeong Park, Junha Jung, Jaewoo Kang