A Controlled Study of Feature-Based Knowledge Distillation Across Student Designs
arXiv:2608. 08294v1 Announce Type: new Abstract: Knowledge distillation trains a smaller student to match the outputs of a larger teacher.
arXiv:2605. 31191v2 Announce Type: replace Abstract: We investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on CIFAR-10.
arXiv:2608. 08294v1 Announce Type: new Abstract: Knowledge distillation trains a smaller student to match the outputs of a larger teacher.
arXiv:2609.22566v1 Announce Type: cross Abstract: Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students. However, distilled students often exhibit subst...
arXiv:2607. 07050v3 Announce Type: replace-cross Abstract: Top-K teacher logits make on-policy distillation tractable, but probability mass is not the same as decision support.
arXiv:2609.13199v1 Announce Type: new Abstract: Knowledge distillation aims to improve the performance of lightweight student models by transferring knowledge from larger and more powerful teacher mo...
arXiv:2606. 25488v1 Announce Type: new Abstract: Knowledge Distillation (KD) is widely used to obtain compact models for efficient inference in resource-constrained environments.
arXiv:2609.39338v1 Announce Type: new Abstract: Knowledge distillation transfers knowledge by encouraging a student to match a teacher's predicted class probabilities. These probabilities express not...
arXiv:2607. 27054v1 Announce Type: new Abstract: Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression.
arXiv:2606. 12171v1 Announce Type: cross Abstract: Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in probability distributions, and mixup enforces them through convex combinations of inputs.
arXiv:2606. 00928v1 Announce Type: cross Abstract: Multiplexed fluorescence microscopy improves tissue segmentation by providing complementary channels including nuclear (DAPI) and membrane (E-cadherin), that together encode richer spatial context than single-channel imaging alone.
The study examines function routing for a 740‑instance healthcare API task using a 1.5B Qwen student and a 20B teacher, comparing eight knowledge‑distillation (KD) variants to supervised cross‑entropy across multiple random seeds. Results show high per‑seed variability (2.8–48.7 pp), with several KD methods exhibiting bimodal collapse—some seeds achieving low accuracy while others train normally—and distinct failure modes such as wrong‑function selection and output‑truncation. Only progressive_kd and rank_kd consistently avoid collapse, and a simple input‑enrichment trick that appeared beneficial in single‑seed tests actually harms performance when re‑tested with multiple seeds.
arXiv:2609.36734v1 Announce Type: new Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implic...
The study investigates what knowledge a student model inherits from its teachers beyond accuracy when using knowledge distillation for encrypted‑traffic classification. By distilling a 101k‑parameter student from two teachers of equal accuracy but different construction, the authors test ten hypotheses over a year of real TLS traffic, finding that unknown‑traffic detection and shortcut reliance can transfer depending on temperature settings and model size, while other abilities do not. The results show that distillation can propagate teacher habits, but some inherited capabilities can also be achieved without a teacher.