arXiv AI By Donggeon Lee, Dooyeon Na, Seungmin Oh, Jongbin Ryu

Layer-wise Curriculum Learning for Efficient LLM Compression

Read the original on arXiv AI →

The paper proposes a layer-wise curriculum learning strategy for compressing large language models (LLMs). By partitioning the model into layer segments and starting training with easier tasks before progressing to harder ones, the method accelerates convergence and stabilizes knowledge transfer from teacher to student models. Additional techniques such as feature caching with multi-threading improve GPU utilization, leading to state‑of‑the‑art compression results and over 50% reductions in memory usage and training time on BERT and GPT‑2, while outperforming other pruning methods on LLaMA‑family and Qwen models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 5

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

arXiv:2608. 03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD).

By Bakbergen Ryskulov, Iker Garc\'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom\'an Or\'us
arXiv Computation and Language
Sep 25

MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression

MILO is a compression framework that reduces the key-value cache memory used in many-shot in-context learning by applying block-wise low-rank compression. It dynamically allocates rank budgets to blocks based on information entropy, preserving important information while aggressively compressing redundant parts. Experiments on Qwen2.5 models show up to a 50% reduction in KV cache memory and a 1.8× throughput improvement with negligible performance loss on classification and reasoning tasks.

By Youpeng Zhao, Tian Tan, Liqian Peng, Jun Wang, Alec Go