arXiv AI By Yiru Yang, Junling Wang, Nishant Kumar Singh, Luohong Wu, Haoran Yan

Logit Distillation on Manifolds: Mapping by Learning

Read the original on arXiv AI →

arXiv:2606. 00771v1 Announce Type: cross Abstract: A simple way to improve the performance of almost any machine learning model is not to train a single but several models with diverse algorithms which will make slightly distinct kinds of predictions and errors on the same data, and thus improve the average predictions and robustness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting

The paper proposes Self‑Distillation Fine‑Tuning (SDFT) as a method to recover performance in Large Language Models that has been degraded by catastrophic forgetting, quantization, or pruning. It shows that SDFT restores model capabilities by aligning the high‑dimensional manifold of the student model’s hidden layers with that of a teacher model, as measured by Centered Kernel Alignment (CKA). The authors provide both empirical evidence of strong correlation between manifold alignment and performance recovery and a theoretical explanation linking generative capability to the structure of these manifolds.

By Chi Liu, Xin Chen, Xu Zhou, Fangbo Tu, Srinivasan Manoharan