The paper proposes Self‑Distillation Fine‑Tuning (SDFT) as a method to recover performance in Large Language Models that has been degraded by catastrophic forgetting, quantization, or pruning. It shows that SDFT restores model capabilities by aligning the high‑dimensional manifold of the student model’s hidden layers with that of a teacher model, as measured by Centered Kernel Alignment (CKA). The authors provide both empirical evidence of strong correlation between manifold alignment and performance recovery and a theoretical explanation linking generative capability to the structure of these manifolds.
By Chi Liu, Xin Chen, Xu Zhou, Fangbo Tu, Srinivasan Manoharan
arXiv:2607. 15919v1 Announce Type: cross Abstract: Data-free knowledge distillation transfers the knowledge encoded in a teacher model to a student model without access to the original training data.
By Mohamed Amine Kina
arXiv:2412. 01282v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks.
By Qianhan Feng, Wenshuo Li, Tong Lin, Xinghao Chen
arXiv:2603. 14830v3 Announce Type: replace Abstract: Dataset distillation, a training-aware data compression technique, has recently attracted increasing attention as an effective tool for mitigating costs of optimization and data storage.
By Yuri Kinoshita, Naoki Nishikawa, Taro Toyoizumi
Large language models (LLMs) achieve strong performance across many tasks, but their high computational cost limits deployment in resource-constrained environments. Knowledge Distillation (KD) offers a practical solution by transferring knowledge from a teacher model of a larger size to a smaller student model.
arXiv:2609.40047v1 Announce Type: cross
Abstract: Gromov-Wasserstein multidimensional scaling (GW-MDS) learns low-dimensional representations from relational data but remains transductive, providing...
By Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante
The paper presents a method for distilling tabular foundation models (TFMs) into lightweight, dataset‑specific students. By using the full labeled training set as teacher context and training students on both observed and synthetic queries, the authors achieve significant performance gains over traditional supervised models on TabArena and TALENT benchmarks. The distilled students also provide substantial inference speedups, reducing the cost of repeated inference.
By Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
The paper investigates how knowledge distillation (KD) applied at intermediate layers of a neural network can affect overfitting and model performance. While traditional KD focuses on the final output, this study explores block‑wise KD across eleven datasets, finding that on standard datasets the last block suffices, but on fine‑grained, data‑scarce settings intermediate supervision significantly improves accuracy. The authors also analyze optimal supervision granularity using attention maps, Centered Kernel Alignment, and Grad‑CAM, and examine teacher‑student fine‑tuning strategies.
By Irene Trigueros-Lorca, Leonardo Concepci\'on, Christian Wagner, Isaac Triguero, Daniel Molina
arXiv:2402. 14035v4 Announce Type: replace-cross Abstract: Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality.
By Zichang Liu, Qingyun Liu, Yuening Li, Liang Liu, Anshumali Shrivastava, Shuchao Bi, Lichan Hong, Ed H. Chi, Zhe Zhao
Cross-modal knowledge distillation (CMKD) studies how a (large) teacher model trained on one type of data (e. g.
arXiv:2607. 23346v1 Announce Type: new Abstract: Modern deep neural networks are potent catalysts for scientific and industrial impact, yet excessive parameter counts impede deployment in low-compute settings such as hospital equipment and energy infrastructure.
By Aditya Dewan, Arjun Yogeswaran, Benjamin Fedoruk
arXiv:2606. 10504v1 Announce Type: new Abstract: Cross-modal knowledge distillation (CMKD) studies how a (large) teacher model trained on one type of data (e.
By Trong Khiem Tran, Anh Duc Chu, Quang Hung Pham, Phi Le Nguyen, Trong Nghia Hoang