arXiv Machine Learning By Jiabin Shen, Guang Chen, Chengjun Mao

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation

Read the original on arXiv Machine Learning →

arXiv:2607. 07050v3 Announce Type: replace-cross Abstract: Top-K teacher logits make on-policy distillation tractable, but probability mass is not the same as decision support.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
1d ago

SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation

arXiv:2608. 14647v1 Announce Type: cross Abstract: Dirty-history rollouts make multi-turn on-policy self-distillation (OPSD) brittle: once a student emits an erroneous intermediate reply, later turns are conditioned on that reply, and uniform distillation can spend loss on tokens that carry little corrective signal.

By Chenyang Jiang, Changhan Huang