arXiv AI By Yixiang Liu, Zhongxing Xu, Zhonghua Wang, Xiaoying Tang

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

Read the original on arXiv AI →

The paper introduces a routing-by-reasoning-need controller for diffusion vision‑language models that dynamically selects early commitment, baseline preservation, or reasoning‑supportive decoding based on trajectory signals such as answer closure, commitment evidence, and representation revision pressure. This training‑free approach avoids a single universal generation length, instead tailoring inference‑time control to the reasoning demands of each question. Experiments on answer‑focused, mixed‑reasoning, and chain‑of‑thought benchmarks show that routed control improves robustness compared to fixed long or short decoding and single‑rule interventions, with benefits not solely due to shorter outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

arXiv:2607. 11436v1 Announce Type: new Abstract: Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning.

By Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang, Zheng Wang, Xing Xu, Jingkuan Song, Yun Zhang, Sirui Da, Heng Tao Shen
arXiv AI
Jul 29

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

arXiv:2607. 25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens.

By Yutong Chen, Shouqian Shi, Xinran Liu, Haochen Wang, Jiaying Wang, Tianxing Xu, Yuanxi Wang, Zirui Ding
arXiv Computer Vision
4d ago

Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning

The paper introduces FlipDir, a training‑free inference‑time technique that mitigates answer flips in vision‑language models by steering hidden states along a low‑rank subspace derived from contrastive image pairs. It employs a margin‑based gate to attenuate steering only during uncertain decoding steps, thereby restoring original predictions while keeping stable ones unchanged. The authors also present VisFlip, a benchmark framework that evaluates models across nine dataset‑variation combinations in scientific reasoning, robot‑scene understanding, and medical VQA, showing that FlipDir consistently outperforms existing methods on recovery and preservation metrics.

By Yeonsung Jung, Joonhyun Jeong, Hoang Pham, Joowon Kim, Yoonsik Park, Viet Dac Lai, Eunho Yang