SPyCE: Skill-Policy Co-evolution for Multimodal Agents
arXiv:2607. 13854v2 Announce Type: replace Abstract: Multimodal agents that think with images iteratively manipulate visual evidence and invoke tools across many steps.
arXiv:2603. 12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestration in open-ended settings.
arXiv:2607. 13854v2 Announce Type: replace Abstract: Multimodal agents that think with images iteratively manipulate visual evidence and invoke tools across many steps.
arXiv:2605. 13527v3 Announce Type: replace Abstract: Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines.
arXiv:2607. 08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing.
arXiv:2608.22963v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decomposition and tool...
Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision.
Beacon is a new agentic visual reasoning model that improves multimodal large language models (MLLMs) by better deciding when to use tools and how to use them. It introduces two key concepts—Mode Adaptiveness, which ensures tools are invoked only when necessary, and Tool Effect, which measures the net benefit of tool use— and trains the model with supervised fine‑tuning and reinforcement learning that rewards necessity-aware decisions and expands capability through expert hints. Across 13 benchmarks, Beacon outperforms other open‑source models, achieving the highest average score and the largest net tool‑gain on diagnostic tests.
OmniHarness is a framework that enables generalizable visual generation by learning symbolic policies from verified executions. It abstracts shared procedures and applicability conditions, allowing these policies to be instantiated, adapted, and composed for new tasks while keeping model parameters fixed. The system uses intermediate verification for refinement, self-directed inquiry to generate practice tasks, and continuous feedback to expand capabilities, achieving strong results on multiple benchmarks and outperforming baselines on Creative tasks.
arXiv:2606. 27974v1 Announce Type: cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowledge.
VISTA is a method for active multimodal agents that internalizes collective visual experience via on‑policy distillation. It turns observations from multiple rollouts of the same input into shared supervision, using Collective Visual Experience Distillation (CVED) to organize observations with context and Heterogeneity‑Aware Policy Improvement (HAPI) to reinforce successful trajectories and guide learning from unsuccessful ones. The approach lets an experience‑conditioned teacher evaluate a student’s partial responses, enabling discoveries from one trajectory to inform others without altering the student’s original history, and achieves superior performance on fine‑grained perception and general reasoning tasks compared to comparable agents.
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-intensive generation.
arXiv:2608. 08557v1 Announce Type: cross Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding.
arXiv:2604.07146v3 Announce Type: replace Abstract: Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for...