arXiv AI

Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold

arXiv AI
Aug 25

Answer First, Reason Later: When Commitment Order Costs Accuracy in Diffusion Language Models

The paper studies how the order in which tokens are committed in masked diffusion language models affects accuracy. It finds that when the final answer is committed before the preceding reasoning (an answer‑first trajectory), accuracy can suffer compared to unrestricted decoding, especially on tasks like GSM8K and MATH‑500. Experiments with controlled token positions show that delaying the answer token can improve performance, indicating that commitment order influences the context and output allocation of the model.

By Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Hwiyeong Lee, Taesup Kim
arXiv Computer Vision
Sep 23

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.

By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao