arXiv AI By Xiaofei Feng

Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation

Read the original on arXiv AI →

The paper investigates correctness‑gated multi‑teacher distillation, comparing a weighted arm to unfiltered distillation across eight experimental arms. While the weighted arm shows modest gains in accuracy (+0.1660) and macro‑F1 (+0.1323) and a reduction in unsafe action rate (−0.4979), it also exhibits lost label functionality, such as zero Refuted recall and over‑assignment of NotEnoughInfo. A subsequent grounding audit was inconclusive, failing to demonstrate a clear improvement in evidence grounding or overall system performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.

By Chencheng Zhu