Discovering and Preserving Category Correlation Knowledge via Adaptive Reciprocal Knowledge Distillation
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.22566v1 Announce Type: cross Abstract: Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students. However, distilled students often exhibit subst...
The paper investigates how knowledge distillation (KD) applied at intermediate layers of a neural network can affect overfitting and model performance. While traditional KD focuses on the final output, this study explores block‑wise KD across eleven datasets, finding that on standard datasets the last block suffices, but on fine‑grained, data‑scarce settings intermediate supervision significantly improves accuracy. The authors also analyze optimal supervision granularity using attention maps, Centered Kernel Alignment, and Grad‑CAM, and examine teacher‑student fine‑tuning strategies.
arXiv:2606. 12171v1 Announce Type: cross Abstract: Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in probability distributions, and mixup enforces them through convex combinations of inputs.
arXiv:2402. 14035v4 Announce Type: replace-cross Abstract: Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality.
arXiv:2606. 03052v1 Announce Type: new Abstract: Knowledge Distillation (KD) is a powerful tool for model compression, yet the precise mechanisms by which student models acquire feature representations remain underexplored.
The paper introduces Teacher Prediction Refinement Distillation (TPRD), a plug‑in module for Detection Transformers that refines teacher predictions before distillation. TPRD corrects degraded positive predictions and suppresses overconfident negatives, while preserving informative dark knowledge through Maximum Dark Knowledge Preservation. Experiments on MS COCO and PASCAL VOC show that these refinements improve the quality of supervision and the resulting student model’s performance.