Hugging Face Trending Papers

When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation

Read the original on Hugging Face Trending Papers →

The paper investigates how the optimal teacher model size for knowledge distillation changes with the amount of training data. It finds that when data is scarce, smaller teachers can outperform larger ones, a phenomenon driven by both score geometry and relational ordering of class predictions. The authors propose DVA, a data‑selection method that uses a small teacher to filter samples by difficulty and maximize diverse relational signals, achieving competitive results without relying on training dynamics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
4d ago

Distillation of Tabular Foundation Models into Efficient Predictors

The paper presents a method for distilling tabular foundation models (TFMs) into lightweight, dataset‑specific students. By using the full labeled training set as teacher context and training students on both observed and synthetic queries, the authors achieve significant performance gains over traditional supervised models on TabArena and TALENT benchmarks. The distilled students also provide substantial inference speedups, reducing the cost of repeated inference.

By Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
arXiv AI
Sep 15

Data-free On-policy Distillation

arXiv:2609.14193v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes...

By Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu, Ruiqi Liu, Qiyong Zhong, Haiyun Guo, Junfeng Fang, Jinqiao Wang
arXiv AI
Sep 7

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

The paper investigates data efficiency and selection in On‑Policy Distillation (OPD) for large language models. It shows that 1‑shot OPD—training on a single example—consistently improves performance, especially when the example is hard, and that longer chain‑of‑thought (CoT) paths drive the gains rather than token entropy. Based on these findings, the authors propose a simple hard‑example selection strategy that, using only eight carefully chosen hard examples, matches the performance of a 17,000‑example baseline across models from 1.5B to 7B parameters.

By Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You