LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
LEGO-OPD introduces a factorized teacher composition for multimodal on‑policy distillation, combining a Language Expert and a Grounding Expert into a single teacher distribution. By treating the language expert as a prior over tokens and the grounding expert as a visual likelihood that updates this prior, the method decouples language reasoning from visual grounding. Adaptive calibration further adjusts the influence of visual evidence at each decoding prefix, preventing over‑ or under‑supervision. Experiments with Qwen3 models demonstrate that LEGO‑OPD outperforms both single‑ and multi‑teacher baselines on multimodal and text‑only reasoning tasks, improving visual perception while preserving language reasoning.
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories w...
OPD‑Aha is a privileged on‑policy distillation method that improves multimodal reasoning by reconstructing the distillation target from the teacher’s isolated visual preference instead of relying on fragile teacher‑student discrepancies. It suppresses continuations that contradict the image, encouraging students to interrupt flawed reasoning with reflection tokens such as "wait" and "actually." This approach leads to consistent improvements across fine‑grained perception and complex multimodal reasoning benchmarks.
arXiv:2606. 05718v1 Announce Type: cross Abstract: On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher.
arXiv:2606. 19120v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target.
Video-OPSD introduces a post‑training framework for Video Large Language Models that leverages privileged visual evidence to enhance on‑policy self‑distillation. The method constructs a self‑teacher conditioned only on annotated evidence frames, while the student processes the full video, allowing the teacher to provide more focused supervision. Additionally, an evidence‑guided token optimization scheme weights distillation based on each token’s reliance on privileged evidence, improving perceptually grounded reasoning. Experiments demonstrate consistent gains over standard OPSD and comparable performance to GRPO with less training time.