VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents
Read the original on arXiv Computer Vision →VISTA is a method for active multimodal agents that internalizes collective visual experience via on‑policy distillation. It turns observations from multiple rollouts of the same input into shared supervision, using Collective Visual Experience Distillation (CVED) to organize observations with context and Heterogeneity‑Aware Policy Improvement (HAPI) to reinforce successful trajectories and guide learning from unsuccessful ones. The approach lets an experience‑conditioned teacher evaluate a student’s partial responses, enabling discoveries from one trajectory to inform others without altering the student’s original history, and achieves superior performance on fine‑grained perception and general reasoning tasks compared to comparable agents.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.