On-Policy Visual Evidence Distillation
arXiv:2609.36838v1 Announce Type: cross Abstract: Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teache...
VISTA is a method for active multimodal agents that internalizes collective visual experience via on‑policy distillation. It turns observations from multiple rollouts of the same input into shared supervision, using Collective Visual Experience Distillation (CVED) to organize observations with context and Heterogeneity‑Aware Policy Improvement (HAPI) to reinforce successful trajectories and guide learning from unsuccessful ones. The approach lets an experience‑conditioned teacher evaluate a student’s partial responses, enabling discoveries from one trajectory to inform others without altering the student’s original history, and achieves superior performance on fine‑grained perception and general reasoning tasks compared to comparable agents.
arXiv:2609.36838v1 Announce Type: cross Abstract: Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teache...
arXiv:2606. 05718v1 Announce Type: cross Abstract: On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher.
arXiv:2609.15683v1 Announce Type: new Abstract: While In-Context Learning (ICL) enables models to adapt from exemplars without parameter updates, multimodal ICL remains largely underexplored, particu...
arXiv:2607. 13854v2 Announce Type: replace Abstract: Multimodal agents that think with images iteratively manipulate visual evidence and invoke tools across many steps.
arXiv:2603. 12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestration in open-ended settings.
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories w...
Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference cost from multi-step diffusion.
arXiv:2606.18974v3 Announce Type: replace Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly a...
OPD‑Aha is a privileged on‑policy distillation method that improves multimodal reasoning by reconstructing the distillation target from the teacher’s isolated visual preference instead of relying on fragile teacher‑student discrepancies. It suppresses continuations that contradict the image, encouraging students to interrupt flawed reasoning with reflection tokens such as "wait" and "actually." This approach leads to consistent improvements across fine‑grained perception and complex multimodal reasoning benchmarks.
VideoGen-Agent is a multimodal agent that uses multitask agentic reinforcement learning to coordinate external tools for video generation. It learns to augment, generate, and verify videos through multi‑turn interactions, guided by prompts and intermediate observations. On the new VABench benchmark, the agent improves base text‑to‑video performance by 19.1 points, and further upgrades to generation tools raise the score to 86.1, with human raters favoring the upgraded configuration in 84.3% of comparisons.
Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision.
arXiv:2609.37923v1 Announce Type: new Abstract: Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce E...