arXiv Computation and Language By Xueqing Wu, Yu-Chi Lin, Kai-Wei Chang, Nanyun Peng

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

Read the original on arXiv Computation and Language →

The paper investigates why post‑training boosts reasoning more than perception in vision‑language models. Using a diagnostic framework with synthetic tasks, it finds a perception‑reasoning asymmetry: supervised fine‑tuning suffers from token imbalance, while reinforcement learning suffers from reward coupling. The authors propose reweighting losses and perception‑aware rewards, achieving up to 18.2‑point and 6.0‑point gains respectively, and show that these methods also improve real‑world visual reasoning by up to 3.3 points.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 17

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment

arXiv:2607. 14682v1 Announce Type: new Abstract: Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge.

By Harikrishnan P M, Goutham Vignesh, Ganesh Parab, Saisubramaniam Gopalakrishnan, Vishal Vaddina, Varun V, Rohit Agrawal