The paper introduces a framework for collaborative memory in multi‑agent vision‑language model (VLM) systems, addressing how agents share and update visual context across distributed perception and reasoning tasks. It outlines a memory hierarchy, cross‑agent sharing protocols, and consistency mechanisms to reconcile differing interpretations and recover missing visual information. The design emphasizes preserving not only raw images or textual summaries but also the dependencies among observations, interpretations, and subsequent reasoning, thereby shaping information flow across agents.
By Huixin Zhang, Shao-Jun Xia, Di Wang, Liangxi Liu, Hainan Xiong, Zihao Wang
Beacon is a new agentic visual reasoning model that improves multimodal large language models (MLLMs) by better deciding when to use tools and how to use them. It introduces two key concepts—Mode Adaptiveness, which ensures tools are invoked only when necessary, and Tool Effect, which measures the net benefit of tool use— and trains the model with supervised fine‑tuning and reinforcement learning that rewards necessity-aware decisions and expands capability through expert hints. Across 13 benchmarks, Beacon outperforms other open‑source models, achieving the highest average score and the largest net tool‑gain on diagnostic tests.
By Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying
The paper introduces Mosaic, a multi-image visual harness that lets large language‑vision models (MLLMs) construct visual intermediates using ten composable image operations. It evaluates five re‑representation settings on existing multi‑image benchmarks and a new grounding‑focused benchmark, MosaicBench, finding that visual re‑representation benefits tasks requiring precise visual evidence more than those dominated by high‑level semantics. MosaicAgent‑8B is trained via reinforcement learning to compose these operations without demonstration trajectories, demonstrating diverse problem‑solving patterns.
By Gengyuan Zhang, Xiao Han, Xinyu Xie, Tong Liu, Volker Tresp
Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when handling ambiguous intentions, logical reasoning, and Out-of-Distribution (OOD) knowledge, existing image models often yield sub-optimal results due to a lack of deep reasoning capabilities and real-time external information.
V‑Retrver is an evidence‑driven retrieval framework that treats universal multimodal retrieval as an agentic reasoning process grounded in visual inspection. It allows multimodal large language models to selectively acquire visual evidence through external tools, alternating between hypothesis generation and targeted visual verification. The approach is trained with a curriculum that blends supervised activation, rejection‑based refinement, and reinforcement learning, achieving an average 23.0% improvement in retrieval accuracy across multiple benchmarks.
By Dongyang Chen, Chaoyang Wang, Dezhao Su, Xi Xiao, Zeyu Zhang, Jing Xiong, Qing Li, Yuzhang Shang, Shichao Kan
arXiv:2607. 14256v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases.
By Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan, Nichole J. Hansen, Bla\v{z} Bratani\v{c}, Nathan L Clement, Shalini Ghosh, Ariel Fuxman