arXiv AI

Transcribe, Then Reason: Two-Pass Decomposition for Multimodal Review

The paper demonstrates that a single-pass multimodal model struggles to produce faithful, comprehensive reviews of long recordings or documents, often omitting a third of the content and embellishing the rest. By splitting the task into two passes—first transcribing the source and then reviewing the transcript—the authors show improved faithfulness and coverage across a diverse set of 21 sources. The benefit is most pronounced for longer or weaker baseline cases, while the approach introduces new failure modes such as space constraints and memory confabulation.

arXiv Machine Learning
1d ago

Min-Cost Flow Routing for Evidence Assembly in Long Multimodal Documents

The paper introduces <flowreader>, a method that casts evidence selection for long multimodal documents as a minimum‑cost flow problem over a multimodal content graph. It uses spectral decomposition to identify latent query‑relevant aspects and allocates a fixed evidence budget proportionally to their spectral energy, ensuring aspect coverage without a language‑model planning call. On the VisDoMBench benchmark with Qwen3‑VL‑32B, <flowreader> achieves the highest macro accuracy (68.9%) and outperforms prior systems by 2.7 points, while using fewer content nodes and reader tokens.

By Ambuj Mehrish, Sebastiano Vascon
arXiv AI
Sep 2

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.

By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
arXiv Computer Vision
Aug 25

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

FOVEA introduces a cache‑friendly, on‑demand visual evidence adaptation for multimodal speculative decoding, enabling a lightweight draft model to dynamically retrieve a bounded subset of visual memory based on a cumulative‑mass rule. The retrieved visual readout is fused with the draft hidden state via a lightweight gated residual correction, avoiding the insertion of visual tokens into the autoregressive context. Experiments on various vision‑language backbones and benchmarks show that FOVEA improves draft acceptance and speeds up end‑to‑end decoding by up to 2.13× compared to traditional autoregressive decoding.

By Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu, Jingxuan Yu, Peng Fu, Zheng Lin, Weiping Wang, Ding Wang
arXiv AI
6d ago

PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

PreviewDiff is a test‑time search method that uses multimodal critics to guide diffusion model sampling. By decoding partial previews at selected denoising checkpoints, scoring them with a multimodal judge, and branching over semantic prompt edits, it allows the generation process to be edited and rerouted before completion. The approach consistently outperforms budget‑matched Best‑of‑N sampling and scalar‑search baselines on image and video benchmarks, with early interventions and wider search yielding the biggest gains.

By Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li, Tomas Pfister, Yale Song
arXiv AI
Aug 19

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
arXiv AI
Aug 11

LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

arXiv:2503. 06211v3 Announce Type: replace-cross Abstract: Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging.

By Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer