arXiv Machine Learning By Quang Hoang Trung, Quang Huu Hieu, Nguyen Van Hoang Phuc, Vo Nguyen Le Duy

ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models

Read the original on arXiv Machine Learning →

The paper introduces Adaptive Local Relational Alignment (ALRA), a logit‑based knowledge distillation method for autoregressive language models that combines student‑generated token proposals with teacher guidance at each prediction position. ALRA dynamically selects the number of candidate tokens based on the teacher’s probability spread, uses Adaptive Local Divergence to match both mass and relative token distributions, and applies Student‑Weighted Pairwise Relational Alignment to focus on high‑probability token pairs. Experiments on The Pile show that 200M‑ and 500M‑parameter students trained with ALRA outperform the best baseline by roughly 1 percentage point and surpass pre‑training without distillation by over 2 percentage points on nine zero‑shot benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 14

Stable On-Policy Distillation through Adaptive Target Reformulation

arXiv:2601. 07155v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models; however, conventional supervised KD often suffers from a distribution mismatch between training and inference.

By Ijun Jang, Jewon Yeom, Juan Yeo, Hyunggyu Lim, Taesup Kim
arXiv AI
Jul 23

When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

arXiv:2607. 19956v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined.

By Dipto Sumit, Ankan Kumar Roy Srizon, Sadia Khair Rodela, Atia Haque Asha, Mourchona Afrin, Niloy Farhan, Farig Sadeque
arXiv AI
2d ago

Multi-LLM Collaborative Alignment via Stackelberg Games

The paper introduces Stackelberg Alignment, a leader‑follower framework that lets a pool of language models collaborate and improve by learning from each other’s responses. An EXP3 bandit leader adaptively selects instructions based on difficulty and discriminability, while the models act as followers, evaluating peers and learning via DPO or GRPO with Elo‑style reputation weighting and opponent matching. Experiments on diverse benchmarks show that this adaptive curriculum outperforms static baselines by up to 12‑25% and improves multi‑LLM evolution.

By Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov