arXiv AI

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

The paper introduces a routing-by-reasoning-need controller for diffusion vision‑language models that dynamically selects early commitment, baseline preservation, or reasoning‑supportive decoding based on trajectory signals such as answer closure, commitment evidence, and representation revision pressure. This training‑free approach avoids a single universal generation length, instead tailoring inference‑time control to the reasoning demands of each question. Experiments on answer‑focused, mixed‑reasoning, and chain‑of‑thought benchmarks show that routed control improves robustness compared to fixed long or short decoding and single‑rule interventions, with benefits not solely due to shorter outputs.

arXiv AI
Jul 14

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

arXiv:2607. 11436v1 Announce Type: new Abstract: Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning.

By Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang, Zheng Wang, Xing Xu, Jingkuan Song, Yun Zhang, Sirui Da, Heng Tao Shen
arXiv AI
Jul 29

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

arXiv:2607. 25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens.

By Yutong Chen, Shouqian Shi, Xinran Liu, Haochen Wang, Jiaying Wang, Tianxing Xu, Yuanxi Wang, Zirui Ding
arXiv Computer Vision
4d ago

Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning

The paper introduces FlipDir, a training‑free inference‑time technique that mitigates answer flips in vision‑language models by steering hidden states along a low‑rank subspace derived from contrastive image pairs. It employs a margin‑based gate to attenuate steering only during uncertain decoding steps, thereby restoring original predictions while keeping stable ones unchanged. The authors also present VisFlip, a benchmark framework that evaluates models across nine dataset‑variation combinations in scientific reasoning, robot‑scene understanding, and medical VQA, showing that FlipDir consistently outperforms existing methods on recovery and preservation metrics.

By Yeonsung Jung, Joonhyun Jeong, Hoang Pham, Joowon Kim, Yoonsik Park, Viet Dac Lai, Eunho Yang
arXiv Computation and Language
4d ago

Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions

The paper investigates how the order of generating explanations—whether a rationale is produced before or after the answer—affects vision‑language reasoning. By conducting controlled experiments on knowledge‑intensive QA, visual entailment, and compositional grounding tasks, the authors show that larger models are required for reliable rationale‑first generation, while answer‑first generation is less susceptible to format errors. The study concludes that explanation ordering, model scale, pre‑training knowledge, fine‑tuning, and task structure jointly influence prediction accuracy and reasoning faithfulness.

By Siting Liang, Luca Rippe, Omar Adjali, Daniel Sonntag
arXiv Computation and Language
Sep 1

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

The paper introduces a framework that combines world models, which generate concrete visual rollouts of possible futures, with multimodal large language models (MLLMs) that perform abstract reasoning. It proposes a controlled concrete reasoning approach and a new training method called Privileged‑Future On‑Policy Self‑Distillation (PF‑OPSD), which uses ground‑truth future videos as privileged teacher context during training while the student model never sees true futures at test time. Experiments on two human‑verified benchmarks, VRQABench and OpenWorldQA, show that PF‑OPSD improves performance by about 10–11% over baselines and enhances robustness to noisy or conflicting rollouts.

By Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen