arXiv Machine Learning By Zihan Li, Feifei Li, Wenhui Que

DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance

Read the original on arXiv Machine Learning →

arXiv:2608. 14644v1 Announce Type: new Abstract: Real-world LLM deployments increasingly rely on runtime-injected prohibitions--enterprise policies, PII redlines, tool boundaries--that vary per request and per tenant.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 11

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

arXiv:2608. 09836v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses.

By Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou
arXiv AI
Jun 9

Trajectory-Refined Distillation

arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.

By Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang
arXiv Machine Learning
Sep 24

Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following

arXiv:2609.27421v1 Announce Type: new Abstract: Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tun...

By Yanzhao Zheng, Yuanqiang Yu, Tianze Xu, Chao Ma, Zhentao Zhang, Jihuai Zhu, Baohua Dong, Hangcheng Zhu, Ruohui Huang