Hugging Face Trending Papers

When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

Read the original on Hugging Face Trending Papers →

Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined. On the BanSum Bangla summarization benchmark, we find that standard KD improves ROUGE-L by only +0.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Jul 23

When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

arXiv:2607. 19956v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined.

By Dipto Sumit, Ankan Kumar Roy Srizon, Sadia Khair Rodela, Atia Haque Asha, Mourchona Afrin, Niloy Farhan, Farig Sadeque
arXiv Computation and Language
Sep 2

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

arXiv:2609.01532v1 Announce Type: new Abstract: Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits...

By Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih