On-Policy Visual Evidence Distillation
arXiv:2609.36838v1 Announce Type: cross Abstract: Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teache...
arXiv:2606. 05718v1 Announce Type: cross Abstract: On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher.
arXiv:2609.36838v1 Announce Type: cross Abstract: Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teache...
Video-OPSD introduces a post‑training framework for Video Large Language Models that leverages privileged visual evidence to enhance on‑policy self‑distillation. The method constructs a self‑teacher conditioned only on annotated evidence frames, while the student processes the full video, allowing the teacher to provide more focused supervision. Additionally, an evidence‑guided token optimization scheme weights distillation based on each token’s reliance on privileged evidence, improving perceptually grounded reasoning. Experiments demonstrate consistent gains over standard OPSD and comparable performance to GRPO with less training time.
arXiv:2609.39120v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods p...
OPD‑Aha is a privileged on‑policy distillation method that improves multimodal reasoning by reconstructing the distillation target from the teacher’s isolated visual preference instead of relying on fragile teacher‑student discrepancies. It suppresses continuations that contradict the image, encouraging students to interrupt flawed reasoning with reflection tokens such as "wait" and "actually." This approach leads to consistent improvements across fine‑grained perception and complex multimodal reasoning benchmarks.
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories w...
arXiv:2608. 05131v1 Announce Type: cross Abstract: On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs).
arXiv:2606. 19120v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target.
arXiv:2606.18974v3 Announce Type: replace Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly a...
arXiv:2608. 14144v1 Announce Type: cross Abstract: Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest.
Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference cost from multi-step diffusion.
LEGO-OPD introduces a factorized teacher composition for multimodal on‑policy distillation, combining a Language Expert and a Grounding Expert into a single teacher distribution. By treating the language expert as a prior over tokens and the grounding expert as a visual likelihood that updates this prior, the method decouples language reasoning from visual grounding. Adaptive calibration further adjusts the influence of visual evidence at each decoding prefix, preventing over‑ or under‑supervision. Experiments with Qwen3 models demonstrate that LEGO‑OPD outperforms both single‑ and multi‑teacher baselines on multimodal and text‑only reasoning tasks, improving visual perception while preserving language reasoning.
arXiv:2610.02117v1 Announce Type: cross Abstract: On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a froze...