arXiv Machine Learning By Jiabin Shen, Guang Chen, Chengjun Mao

Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation

Read the original on arXiv Machine Learning →

arXiv:2607. 07050v1 Announce Type: cross Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

The paper investigates how teacher signals influence parameter updates in Multi‑Teacher On‑Policy Distillation (MOPD) by analyzing Qwen3‑1.7B and SmolLM3‑3B. It shows that loss averaging, Adam’s first‑moment bias, BF16 rounding, and the choice of averaging rule all shape the gradients and ultimately affect task performance. The study quantifies these effects, revealing, for example, that token‑averaging favors longer responses and that BF16 rounding masks most weight changes.

By Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang, Zhanyang Jin, Yihang Sun, Jiaxuan You