arXiv Machine Learning By Jiabin Shen, Guang Chen, Chengjun Mao

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation

Read the original on arXiv Machine Learning →

arXiv:2607. 07050v3 Announce Type: replace-cross Abstract: Top-K teacher logits make on-policy distillation tractable, but probability mass is not the same as decision support.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 4

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

The paper introduces Teacher-Gated On-Policy Distillation (TGOPD), a method that verifies teacher reliability at the prompt level before applying dense supervision in on-policy distillation. TGOPD uses verifier-scored teacher probes to decide whether to route a prompt to dense OPD or to a verifier-grounded alternative. Experiments on 4B and 35B models across mathematics, code, and instruction tasks show TGOPD outperforms vanilla OPD and improves teacher GPU utilization from 9.8% to 78.9% in a 4B single-domain run.

By Zhiwei Zhang, Zechen Sun, Fei Zhao, Kang Peng, Bin Liang, Huayu Deng, Yao Hu, Kam-Fai Wong, Mu Chuan
arXiv Machine Learning
Sep 22

Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy

The paper presents a causal analysis of a compressed VLA policy that performs well in offline tests but fails in closed‑loop execution on a simulated pick‑and‑place task. An 8‑layer distillation of Octo‑Base retains most parameters and passes all offline metrics, yet collapses during deployment, with early stages degrading gradually and final transport failing entirely. The failure is traced to a negative, late‑heavy residual in the action trace, and standard remedies (continued training, offline data, command‑level compensation, clamping) do not restore performance; only a minimal‑pair intervention that mixes deployment‑distribution rollouts with teacher data restores parity with the teacher. whyItMatters":"The study demonstrates that offline validation metrics alone are insufficient to guarantee closed‑loop success for compressed policies, highlighting the need for targeted deployment‑time testing and interventions."

By Fengze Jia (The Ohio State University)