arXiv Machine Learning By Yongkang Zhou, Xiang Xia, Cheng Yan, Fan Xu, Wuyang Zhang

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

Read the original on arXiv Machine Learning →

arXiv:2608. 01821v1 Announce Type: cross Abstract: Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial recurring inference cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Aug 25

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

FOVEA introduces a cache‑friendly, on‑demand visual evidence adaptation for multimodal speculative decoding, enabling a lightweight draft model to dynamically retrieve a bounded subset of visual memory based on a cumulative‑mass rule. The retrieved visual readout is fused with the draft hidden state via a lightweight gated residual correction, avoiding the insertion of visual tokens into the autoregressive context. Experiments on various vision‑language backbones and benchmarks show that FOVEA improves draft acceptance and speeds up end‑to‑end decoding by up to 2.13× compared to traditional autoregressive decoding.

By Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu, Jingxuan Yu, Peng Fu, Zheng Lin, Weiping Wang, Ding Wang
arXiv AI
3d ago

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

GroundAnything is a 4‑B parameter grounding foundation model that combines autoregressive and diffusion approaches to achieve fast parallel decoding while maintaining precise visual grounding. By treating grounding as visual evidence extraction and using blockwise denoising, it allows spatial hypotheses to be generated in parallel and refined iteratively. The model outperforms existing state‑of‑the‑art methods on 30 grounding benchmarks, achieving 72.42% accuracy with its autoregressive variant and 61.75% with entropy‑guided decoding, while also offering significant speedups through optional self‑speculative decoding.

By Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou, Ping Luo, Shiyu Huang
arXiv AI
Aug 11

Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression

arXiv:2608. 09176v1 Announce Type: cross Abstract: Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost.

By Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu, Kangning Cui, Xilu Wang