arXiv:2512. 21002v3 Announce Type: replace-cross Abstract: Distilling the capabilities from a large reasoning model (LRM) to a smaller student model often involves training on substantial amounts of reasoning data.
By Wei-Rui Chen, Vignesh Kothapalli, Ata Fatahibaarzi, Hejian Sang, Shao Tang, Qingquan Song, Zhipeng Wang, Muhammad Abdul-Mageed
The paper studies how to compress Chain-of-Thought (CoT) reasoning traces for smaller models. It examines three compression dimensions—importance criterion, restructuring level, and compression budget—across Math and General domains and Long/Short CoT regimes. Findings show that step-level pruning works best for shared reasoning backbones, token-level pruning needs symbol-aware signals, domain-specific restructuring effects differ, and training-time compression may not reduce inference cost, especially for Long-CoT students.
By Siyang Lyu, Xinghao Chen, Zhijing Sun, Tong Liu, Dawei Zhu, Xiaoyu Shen
arXiv:2606. 31048v1 Announce Type: cross Abstract: This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.
By Gaurab Baral, Aaditya Khanal, Yangyang Tao, Junxiu Zhou
arXiv:2607.22629v3 Announce Type: replace
Abstract: Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate t...
By Durgesh Kalwar, Vardhan Palod, Jaya Adithya Pavuluri, Subbarao Kambhampati
arXiv:2608. 03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD).
By Bakbergen Ryskulov, Iker Garc\'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom\'an Or\'us
The paper investigates whether distilling reasoning traces from large teacher models is worth the extra compute compared to standard instruction fine‑tuning (IFT). By generating paired IFT and reasoning outputs from the same teacher and training student models at five scales, the authors find that, at matched FLOPs, IFT generally lies on or near the Pareto frontier across most configurations. Reasoning traces only reach the frontier on open‑ended tasks for models 7B and larger, and a curriculum mixing 25–50% reasoning data with IFT can capture most of the accuracy benefit at a lower compute cost.
By Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El Haddad, C\'eline Hudelot, Pierre Colombo