arXiv AI By Haixu Ma, Saad Lahrichi, Weiwei Li, Kevin Han, Weiqiang Wu, Peggy Yang, Dongzhuo Li, Ruiyi Li, Serena Li, Gedi Zhou, Mingze Gao, Abhishek Kumar, Xiangjun Fan, Lizhu Zhang

Persistent Negatives for Adversarial Black-Box On-Policy Distillation

Read the original on arXiv AI →

The paper introduces Persistent Negative Adversarial Distillation, a method that improves black-box on-policy distillation by maintaining a live pool of historical teacher–student comparisons to stabilize the discriminator’s negative distribution. By anchoring the discriminator with these persistent negatives, the approach reduces reward-estimation error and yields smoother, higher-performing student policies across multiple benchmarks. The study demonstrates that the choice of negative samples is a critical design factor in effective black-box distillation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 24

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

The paper introduces Preference‑Based Self‑Distillation (PBSD), a new on‑policy self‑distillation method that replaces traditional KL matching with a reward‑regularized objective. PBSD derives a reward‑reweighted teacher distribution, optimizing preference gaps between teacher and student samples while keeping on‑policy sampling. Experiments on mathematical reasoning and tool‑use tasks show PBSD achieves stronger average performance, improved training stability, and maintains token efficiency compared to prior self‑distillation baselines.

By Xin Yu, Liuchen Liao, Yiwen Zhang, Yingchen Yu, Lingzhou Xue, Qinzhen Guo
arXiv AI
Jul 7

Reward-Gated On-Policy Distillation

arXiv:2607. 04037v1 Announce Type: cross Abstract: On-policy distillation is a powerful way to transfer reasoning ability from a strong teacher to a smaller student: the student samples trajectories from its own policy, and the teacher provides dense token-level supervision on the states the student actually visits.

By Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi, Chanakya Ekbote, Sujay Sanghavi, Aleksandar Bojchevski
arXiv Machine Learning
Sep 1

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

The paper investigates on‑policy distillation (OPD), showing that teacher supervision during OPD contains significant noise that grows with teacher size, yet the student policy remains largely unaffected by this noise. It finds that OPD’s gains stem mainly from suppressing low‑log‑probability tokens, a process that can be replicated without a teacher. Building on this insight, the authors propose On‑Policy Self‑Adaptation (OPSA), a supervision‑free method that uses entropy‑adaptive negative advantages to improve performance on several benchmarks, outperforming both the base model and OPD.

By Yi Ding, Ruqi Zhang
arXiv AI
Jul 28

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

arXiv:2607. 24731v1 Announce Type: cross Abstract: On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood.

By Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang