arXiv AI

Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering

arXiv:2608. 04124v1 Announce Type: cross Abstract: Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time.

arXiv AI
Sep 10

Reason Through the Latent! Making Latent Visual Reasoning Necessary

The paper introduces Causal Visual Recurrent Reasoning (CVRR), a method that forces multimodal models to rely on latent visual states by using recurrent computation as the sole image‑conditioned path to prediction. CVRR initializes recurrence from the question hidden state after a pretrained vision‑language model has processed the image, repeatedly updates this state while re‑reading the same visual evidence, and removes all other visual traces before decoding. Experiments on several benchmarks show that CVRR maintains strong performance while other latent reasoners lose visual competence, and causal interventions confirm that predictions depend on the recurrent visual trajectory.

By Suhyeong Park, Junha Jung, Jaewoo Kang
arXiv Machine Learning
Jun 2

Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering

arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.

By Ali Alavi
arXiv AI
Jul 21

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.

By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
arXiv AI
Aug 25

VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting

VisionCoach is an input‑adaptive reinforcement learning framework that enhances spatio‑temporal grounding in video reasoning by using visual prompting during training. The system selectively applies visual prompts to challenging inputs, amplifying question‑relevant evidence and suppressing distractors, and then internalizes these improvements through self‑distillation so that inference can be performed on raw videos without prompts. Experiments on multiple benchmarks (V‑STAR, VideoMME, World‑Sense, VideoMMMU, PerceptionTest, and Charades‑STA) show that VisionCoach achieves state‑of‑the‑art performance while maintaining a single efficient inference pathway.

By Daeun Lee, Shoubin Yu, Yue Zhang, Mohit Bansal
arXiv Computer Vision
Aug 28

Video-FLAIR: Not Whether to Reason, But How

Video-FLAIR is a training framework that teaches a model to choose the most suitable reasoning mode—perceptual, compositional, or deliberative—for each multimodal query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt, and a composite reward system selects the best response based on correctness, grounding, and cost, discouraging unsupported deliberation. This adaptive approach improves accuracy on benchmarks such as MathVista, Video-Holmes, and Video-MMMU while dramatically reducing token usage compared to always-thinking baselines.

By Yogesh Kulkarni, Pooyan Fazli
arXiv AI
Jun 24

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

arXiv:2606. 24477v1 Announce Type: cross Abstract: Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA).

By Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li, Zejun MA, Chao Zhang
arXiv Machine Learning
Jul 7

Incentivizing Vision Language Models to Search for Long Video Question Answering

arXiv:2607. 02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process.

By Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep P. Chinchali