ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models
Read the original on arXiv Machine Learning →The paper introduces Adaptive Local Relational Alignment (ALRA), a logit‑based knowledge distillation method for autoregressive language models that combines student‑generated token proposals with teacher guidance at each prediction position. ALRA dynamically selects the number of candidate tokens based on the teacher’s probability spread, uses Adaptive Local Divergence to match both mass and relative token distributions, and applies Student‑Weighted Pairwise Relational Alignment to focus on high‑probability token pairs. Experiments on The Pile show that 200M‑ and 500M‑parameter students trained with ALRA outperform the best baseline by roughly 1 percentage point and surpass pre‑training without distillation by over 2 percentage points on nine zero‑shot benchmarks.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.