arXiv:2601. 01484v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher network to a typically smaller student model, often by leveraging soft probabilistic outputs.
By Itai Morad, Nir Shlezinger, Yonina C. Eldar
arXiv:2609.36734v1 Announce Type: new
Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implic...
By Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty
arXiv:2607. 09692v1 Announce Type: new Abstract: Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations.
By Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, Sewon Min
The paper investigates why self‑distillation can sometimes worsen the reasoning abilities of large language models (LLMs). It finds that the process suppresses the model’s epistemic verbalization—its expression of uncertainty—leading to shorter but less accurate responses in mathematical reasoning tasks. Experiments on several LLMs show performance drops of up to 40%, especially on out‑of‑distribution problems where uncertainty expression is beneficial.
By Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang
The paper presents a method for distilling tabular foundation models (TFMs) into lightweight, dataset‑specific students. By using the full labeled training set as teacher context and training students on both observed and synthetic queries, the authors achieve significant performance gains over traditional supervised models on TabArena and TALENT benchmarks. The distilled students also provide substantial inference speedups, reducing the cost of repeated inference.
By Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
arXiv:2602. 01477v2 Announce Type: replace-cross Abstract: Evidential Deep Learning (EDL) is a popular framework for uncertainty-aware classification that models predictive uncertainty via Dirichlet distributions parameterized by neural networks.
By Pietro Carlotti, Nevena Gligi\'c, Arya Farahi