arXiv Machine Learning

Student Capacity Moderates Knowledge Distillation Effectiveness: A Systematic Study Across ResNet Teacher-Student Pairs on CIFAR-10

arXiv:2605. 31191v2 Announce Type: replace Abstract: We investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on CIFAR-10.

arXiv Machine Learning
Jun 2

Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models

arXiv:2606. 00928v1 Announce Type: cross Abstract: Multiplexed fluorescence microscopy improves tissue segmentation by providing complementary channels including nuclear (DAPI) and membrane (E-cadherin), that together encode richer spatial context than single-channel imaging alone.

By Sakib Mohammad, Jarin Ritu, Md Sakhawat Hossain
arXiv Computation and Language
Aug 31

Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation

The study examines function routing for a 740‑instance healthcare API task using a 1.5B Qwen student and a 20B teacher, comparing eight knowledge‑distillation (KD) variants to supervised cross‑entropy across multiple random seeds. Results show high per‑seed variability (2.8–48.7 pp), with several KD methods exhibiting bimodal collapse—some seeds achieving low accuracy while others train normally—and distinct failure modes such as wrong‑function selection and output‑truncation. Only progressive_kd and rank_kd consistently avoid collapse, and a simple input‑enrichment trick that appeared beneficial in single‑seed tests actually harms performance when re‑tested with multiple seeds.

By Dipto Sumit, Sakib Ul Haque, Farig Sadeque
arXiv Machine Learning
5d ago

Unknown-Traffic Detection, Calibration and Shortcut Reliance in Distilled Encrypted-Traffic Classifiers over One Year

The study investigates what knowledge a student model inherits from its teachers beyond accuracy when using knowledge distillation for encrypted‑traffic classification. By distilling a 101k‑parameter student from two teachers of equal accuracy but different construction, the authors test ten hypotheses over a year of real TLS traffic, finding that unknown‑traffic detection and shortcut reliance can transfer depending on temperature settings and model size, while other abilities do not. The results show that distillation can propagate teacher habits, but some inherited capabilities can also be achieved without a teacher.

By Mahmoud Abbasi