arXiv Machine Learning By Maxime Griot, Paul Steven Scotti, Tanishq Mathew Abraham

Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation

Read the original on arXiv Machine Learning →

arXiv:2606. 05988v1 Announce Type: new Abstract: Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 2

When Compression Helps and When It Hurts: Condition-Aware Analysis of Chain-of-Thought Distillation

The paper studies how to compress Chain-of-Thought (CoT) reasoning traces for smaller models. It examines three compression dimensions—importance criterion, restructuring level, and compression budget—across Math and General domains and Long/Short CoT regimes. Findings show that step-level pruning works best for shared reasoning backbones, token-level pruning needs symbol-aware signals, domain-specific restructuring effects differ, and training-time compression may not reduce inference cost, especially for Long-CoT students.

By Siyang Lyu, Xinghao Chen, Zhijing Sun, Tong Liu, Dawei Zhu, Xiaoyu Shen
arXiv AI
Aug 5

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

arXiv:2608. 03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD).

By Bakbergen Ryskulov, Iker Garc\'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom\'an Or\'us
arXiv Computation and Language
Aug 24

Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation

The paper investigates whether distilling reasoning traces from large teacher models is worth the extra compute compared to standard instruction fine‑tuning (IFT). By generating paired IFT and reasoning outputs from the same teacher and training student models at five scales, the authors find that, at matched FLOPs, IFT generally lies on or near the Pareto frontier across most configurations. Reasoning traces only reach the frontier on open‑ended tasks for models 7B and larger, and a curriculum mixing 25–50% reasoning data with IFT can capture most of the accuracy benefit at a lower compute cost.

By Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El Haddad, C\'eline Hudelot, Pierre Colombo