arXiv:2506.02294v4 Announce Type: replace
Abstract: Large foundation models trained on extensive datasets demonstrate strong zero-shot capabilities in various domains. Knowledge distillation has beco...
By Niclas Popp, Kevin Alexander Laube, Matthias Hein, Lukas Schott
arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.
By Trisha Mittal, Akshay Mehra, Joshua Kimball
The paper investigates how knowledge distillation (KD) applied at intermediate layers of a neural network can affect overfitting and model performance. While traditional KD focuses on the final output, this study explores block‑wise KD across eleven datasets, finding that on standard datasets the last block suffices, but on fine‑grained, data‑scarce settings intermediate supervision significantly improves accuracy. The authors also analyze optimal supervision granularity using attention maps, Centered Kernel Alignment, and Grad‑CAM, and examine teacher‑student fine‑tuning strategies.
By Irene Trigueros-Lorca, Leonardo Concepci\'on, Christian Wagner, Isaac Triguero, Daniel Molina
Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher is an effective remedy that leaves the deployed model unchanged. General-purpose feature distillation, however, transfers little in this setting.
arXiv:2608. 09091v1 Announce Type: cross Abstract: Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO .
By Jing Ning, James D. Braza
arXiv:2607. 14703v1 Announce Type: cross Abstract: Multiple instance learning (MIL) has become the main paradigm for whole-slide image (WSI) analysis in computational pathology.
By Mingxi Fu, Jiawen Li, Renao Yan, Jiali Hu, Qiehe Sun, Tian Guan, Yonghong He
arXiv:2606. 25488v1 Announce Type: new Abstract: Knowledge Distillation (KD) is widely used to obtain compact models for efficient inference in resource-constrained environments.
By Yifan Wu, Yiqi Wang, Xichen Ye, Wenjing Yan, Xiaoqiang Li, Cheng Jin, Xiangyu Yue, Weizhong Zhang
arXiv:2607. 09100v1 Announce Type: cross Abstract: The rapid growth of image data has produced large-scale datasets, raising concerns about the time and memory costs of model training.
By Pedro Rocha Dantas, Lucas Pascotti Valem
The paper introduces IDeaL, a data‑free multi‑teacher distillation technique that generates teacher‑specific, improved samples using decorrelation losses at patch and image levels. By tailoring noise to each teacher, IDeaL produces strong student models that capture complementary teacher information and achieve results close to those distilled from real images. Experiments demonstrate that with only 1,000 images, students trained on IDeaL samples match or exceed the performance of students distilled from a 1,000‑image subset of ImageNet.
By Feyza Yavuz, Mert B\"ulent Sar{\i}y{\i}ld{\i}z, Diane Larlus
The paper introduces TALON, a Task‑Adaptive LoRA‑Teacher framework for Few‑Shot Class‑Incremental Learning. TALON assigns a dedicated LoRA‑Teacher to each incremental task, then distills the frozen teachers into a single LoRA‑Student via Ensemble Knowledge Transfer, using a semantic‑guided weighting scheme to reduce forgetting and overfitting. Experiments on four FSCIL benchmarks show that TALON matches or surpasses state‑of‑the‑art accuracy while using up to 33× fewer deployment parameters and cutting inference time by 41.7%.
By Hongwei Zhao (School of Computer Science,Engineering, Beihang University), Rui Liu (School of Computer Science,Engineering, Beihang University), Yansong Liu (School of Computer Science,Engineering, Beihang University), Zhiyuan Zou (School of Computer Science,Engineering, Beihang University), Yong Chen (School of Computer Science, Beijing University of Posts,Telecommunications)
arXiv:2604. 05634v2 Announce Type: replace Abstract: Machine unlearning (MU) has become a critical technique for GenAI models' safe and compliant operation.
By Zhiyong Ma, Zhitao Deng, Huan Tang, Jialin Chen, Zhijun Zheng, Zhengping Li, Qingyuan Chuai
arXiv:2606. 23897v1 Announce Type: cross Abstract: Prompt distillation compresses large vision-language models (VLMs) such as CLIP into lightweight student models by matching teacher predictions on unlabeled domain images.
By Ahmad Algadhi, Ahmed Alzuhair, Omar Alkhulaif, Muzammil Behzad