Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation
Read the original on arXiv Computation and Language →The study examines function routing for a 740‑instance healthcare API task using a 1.5B Qwen student and a 20B teacher, comparing eight knowledge‑distillation (KD) variants to supervised cross‑entropy across multiple random seeds. Results show high per‑seed variability (2.8–48.7 pp), with several KD methods exhibiting bimodal collapse—some seeds achieving low accuracy while others train normally—and distinct failure modes such as wrong‑function selection and output‑truncation. Only progressive_kd and rank_kd consistently avoid collapse, and a simple input‑enrichment trick that appeared beneficial in single‑seed tests actually harms performance when re‑tested with multiple seeds.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.