On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
Read the original on arXiv Computation and Language →The paper investigates why post‑training boosts reasoning more than perception in vision‑language models. Using a diagnostic framework with synthetic tasks, it finds a perception‑reasoning asymmetry: supervised fine‑tuning suffers from token imbalance, while reinforcement learning suffers from reward coupling. The authors propose reweighting losses and perception‑aware rewards, achieving up to 18.2‑point and 6.0‑point gains respectively, and show that these methods also improve real‑world visual reasoning by up to 3.3 points.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.