arXiv AI By Bojie Li, Noah Shi

Transcribe, Then Reason: Two-Pass Decomposition for Multimodal Review

Read the original on arXiv AI →

The paper demonstrates that a single-pass multimodal model struggles to produce faithful, comprehensive reviews of long recordings or documents, often omitting a third of the content and embellishing the rest. By splitting the task into two passes—first transcribing the source and then reviewing the transcript—the authors show improved faithfulness and coverage across a diverse set of 21 sources. The benefit is most pronounced for longer or weaker baseline cases, while the approach introduces new failure modes such as space constraints and memory confabulation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

Min-Cost Flow Routing for Evidence Assembly in Long Multimodal Documents

The paper introduces <flowreader>, a method that casts evidence selection for long multimodal documents as a minimum‑cost flow problem over a multimodal content graph. It uses spectral decomposition to identify latent query‑relevant aspects and allocates a fixed evidence budget proportionally to their spectral energy, ensuring aspect coverage without a language‑model planning call. On the VisDoMBench benchmark with Qwen3‑VL‑32B, <flowreader> achieves the highest macro accuracy (68.9%) and outperforms prior systems by 2.7 points, while using fewer content nodes and reader tokens.

By Ambuj Mehrish, Sebastiano Vascon
arXiv AI
Sep 2

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.

By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
arXiv Computer Vision
Aug 25

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

FOVEA introduces a cache‑friendly, on‑demand visual evidence adaptation for multimodal speculative decoding, enabling a lightweight draft model to dynamically retrieve a bounded subset of visual memory based on a cumulative‑mass rule. The retrieved visual readout is fused with the draft hidden state via a lightweight gated residual correction, avoiding the insertion of visual tokens into the autoregressive context. Experiments on various vision‑language backbones and benchmarks show that FOVEA improves draft acceptance and speeds up end‑to‑end decoding by up to 2.13× compared to traditional autoregressive decoding.

By Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu, Jingxuan Yu, Peng Fu, Zheng Lin, Weiping Wang, Ding Wang