arXiv Computer Vision By Yicheng Xue, Han Wu, Jufeng Yang, Minjing Dong, Xinghao Chen, Hanting Chen, Jianyuan Guo

Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning

Read the original on arXiv Computer Vision →

The paper introduces ViMoD, a lightweight framework that dynamically manages visual memory for multimodal large language models. ViMoD uses Deformable Aggregation of Region-wise Tokens (DART) to create content‑adaptive coarse representations linked to recoverable fine tokens, and Temporal Routing for Adaptive Contextual Evidence (TRACE) to anticipate and adjust evidence needs during reasoning. On the Qwen3‑VL‑4B model, ViMoD achieves superior performance across eight reasoning benchmarks with only a 20% visual token budget and minimal additional parameters.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 28

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

The paper introduces LIRSeg, a method that replaces explicit Chain-of-Thought reasoning in multimodal large language models with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages—spatial alignment and GRPO—while employing extreme-advantage sampling, decoupled exploration-stability updates, and latent diversity amplification to enhance token informativeness. Experiments show that LIRSeg improves segmentation accuracy and reasoning efficiency, achieving significant gIoU gains over the VisionReasoner baseline and reducing reasoning tokens by about 16×.

By Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan