Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 14099v1 Announce Type: cross Abstract: Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure.
The paper studies how the order in which tokens are committed in masked diffusion language models affects accuracy. It finds that when the final answer is committed before the preceding reasoning (an answer‑first trajectory), accuracy can suffer compared to unrestricted decoding, especially on tasks like GSM8K and MATH‑500. Experiments with controlled token positions show that delaying the answer token can improve performance, indicating that commitment order influences the context and output allocation of the model.
arXiv:2607. 16165v1 Announce Type: cross Abstract: Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot.
arXiv:2607. 15565v1 Announce Type: cross Abstract: Where should the question go in a vision-language model (VLM) prompt: before the image or after it?
arXiv:2606. 14620v1 Announce Type: new Abstract: Open diffusion language models are marketed as parallel, non-autoregressive decoders, yet the order in which a shipped checkpoint actually commits its tokens is almost never measured.
arXiv:2607. 12304v1 Announce Type: cross Abstract: A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions.