arXiv Computation and Language By Dipto Sumit, Sakib Ul Haque, Farig Sadeque

Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation

Read the original on arXiv Computation and Language →

The study examines function routing for a 740‑instance healthcare API task using a 1.5B Qwen student and a 20B teacher, comparing eight knowledge‑distillation (KD) variants to supervised cross‑entropy across multiple random seeds. Results show high per‑seed variability (2.8–48.7 pp), with several KD methods exhibiting bimodal collapse—some seeds achieving low accuracy while others train normally—and distinct failure modes such as wrong‑function selection and output‑truncation. Only progressive_kd and rank_kd consistently avoid collapse, and a simple input‑enrichment trick that appeared beneficial in single‑seed tests actually harms performance when re‑tested with multiple seeds.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 23

When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

arXiv:2607. 19956v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined.

By Dipto Sumit, Ankan Kumar Roy Srizon, Sadia Khair Rodela, Atia Haque Asha, Mourchona Afrin, Niloy Farhan, Farig Sadeque
arXiv AI
Aug 18

SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation

arXiv:2608. 14647v1 Announce Type: cross Abstract: Dirty-history rollouts make multi-turn on-policy self-distillation (OPSD) brittle: once a student emits an erroneous intermediate reply, later turns are conditioned on that reply, and uniform distillation can spend loss on tokens that carry little corrective signal.

By Chenyang Jiang, Changhan Huang