arXiv:2601. 01484v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher network to a typically smaller student model, often by leveraging soft probabilistic outputs.
By Itai Morad, Nir Shlezinger, Yonina C. Eldar
arXiv:2609.36734v1 Announce Type: new
Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implic...
By Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty
arXiv:2607. 09692v1 Announce Type: new Abstract: Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations.
By Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, Sewon Min
The paper investigates why self‑distillation can sometimes worsen the reasoning abilities of large language models (LLMs). It finds that the process suppresses the model’s epistemic verbalization—its expression of uncertainty—leading to shorter but less accurate responses in mathematical reasoning tasks. Experiments on several LLMs show performance drops of up to 40%, especially on out‑of‑distribution problems where uncertainty expression is beneficial.
By Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang
The paper presents a method for distilling tabular foundation models (TFMs) into lightweight, dataset‑specific students. By using the full labeled training set as teacher context and training students on both observed and synthetic queries, the authors achieve significant performance gains over traditional supervised models on TabArena and TALENT benchmarks. The distilled students also provide substantial inference speedups, reducing the cost of repeated inference.
By Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
arXiv:2602. 01477v2 Announce Type: replace-cross Abstract: Evidential Deep Learning (EDL) is a popular framework for uncertainty-aware classification that models predictive uncertainty via Dirichlet distributions parameterized by neural networks.
By Pietro Carlotti, Nevena Gligi\'c, Arya Farahi
arXiv:2402. 14035v4 Announce Type: replace-cross Abstract: Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality.
By Zichang Liu, Qingyun Liu, Yuening Li, Liang Liu, Anshumali Shrivastava, Shuchao Bi, Lichan Hong, Ed H. Chi, Zhe Zhao
arXiv:2609.38342v1 Announce Type: new
Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....
By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv:2609.13199v1 Announce Type: new
Abstract: Knowledge distillation aims to improve the performance of lightweight student models by transferring knowledge from larger and more powerful teacher mo...
By Dawen Jiang, Zhishu Shen, Zeyu Liu, Tiehua Zhang
arXiv:2602. 08142v2 Announce Type: replace Abstract: Machine learning applications require fast and reliable per-sample uncertainty estimation.
By H. Martin Gillis, Isaac Xu, Thomas Trappenberg
arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.
By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
The paper introduces IDeaL, a data‑free multi‑teacher distillation technique that generates teacher‑specific, improved samples using decorrelation losses at patch and image levels. By tailoring noise to each teacher, IDeaL produces strong student models that capture complementary teacher information and achieve results close to those distilled from real images. Experiments demonstrate that with only 1,000 images, students trained on IDeaL samples match or exceed the performance of students distilled from a 1,000‑image subset of ImageNet.
By Feyza Yavuz, Mert B\"ulent Sar{\i}y{\i}ld{\i}z, Diane Larlus