arXiv AI By Huipeng Huang, Hongxin Wei

Tail-Aware Top-$k$ On-Policy Distillation

Read the original on arXiv AI →

arXiv:2608. 14728v1 Announce Type: cross Abstract: On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next-token distribution with the teacher's along its own trajectories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.