arXiv:2512. 21002v3 Announce Type: replace-cross Abstract: Distilling the capabilities from a large reasoning model (LRM) to a smaller student model often involves training on substantial amounts of reasoning data.
By Wei-Rui Chen, Vignesh Kothapalli, Ata Fatahibaarzi, Hejian Sang, Shao Tang, Qingquan Song, Zhipeng Wang, Muhammad Abdul-Mageed
The paper studies how to compress Chain-of-Thought (CoT) reasoning traces for smaller models. It examines three compression dimensions—importance criterion, restructuring level, and compression budget—across Math and General domains and Long/Short CoT regimes. Findings show that step-level pruning works best for shared reasoning backbones, token-level pruning needs symbol-aware signals, domain-specific restructuring effects differ, and training-time compression may not reduce inference cost, especially for Long-CoT students.
By Siyang Lyu, Xinghao Chen, Zhijing Sun, Tong Liu, Dawei Zhu, Xiaoyu Shen
arXiv:2606. 31048v1 Announce Type: cross Abstract: This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.
By Gaurab Baral, Aaditya Khanal, Yangyang Tao, Junxiu Zhou
arXiv:2607.22629v3 Announce Type: replace
Abstract: Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate t...
By Durgesh Kalwar, Vardhan Palod, Jaya Adithya Pavuluri, Subbarao Kambhampati
arXiv:2608. 03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD).
By Bakbergen Ryskulov, Iker Garc\'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom\'an Or\'us
The paper investigates whether distilling reasoning traces from large teacher models is worth the extra compute compared to standard instruction fine‑tuning (IFT). By generating paired IFT and reasoning outputs from the same teacher and training student models at five scales, the authors find that, at matched FLOPs, IFT generally lies on or near the Pareto frontier across most configurations. Reasoning traces only reach the frontier on open‑ended tasks for models 7B and larger, and a curriculum mixing 25–50% reasoning data with IFT can capture most of the accuracy benefit at a lower compute cost.
By Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El Haddad, C\'eline Hudelot, Pierre Colombo
arXiv:2609.26708v1 Announce Type: new
Abstract: Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathema...
By Yuanteng Chen, Zhilei Liu, Peisong Wang, Yuantian Shao, Chuangyi Li, Weining Wang, Shuang Qiu, Gang Li, Jing Liu, Jian Cheng
arXiv:2606. 21994v2 Announce Type: replace Abstract: On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories.
By Qingfei Zhao, Huan Song, Shuyu Tian, Jiawei Shao, Xuelong Li
arXiv:2608. 14277v1 Announce Type: cross Abstract: On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including tokenizer mismatch, teacher-student distribution mismatch, response length explosion, and training instability.
By Haonan He, Haodi Lei, Yun Luo, Haoran Zhang, Shunkai Zhang, Yizhuo Li, Shengji Tang, Zhilin Wang, Runzhe Zhan, Lei Bai, Ganqu Cui, Fangchen Yu, Yafu Li, Peng Ye, Ning Ding, Yu Cheng
arXiv:2609.37066v1 Announce Type: cross
Abstract: Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed...
By Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu, Said Mammar, Pascal Bouvry
arXiv:2609.36246v1 Announce Type: new
Abstract: We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressivel...
By Haojin Wang, Dylan Zhang, Huaibo Chen, Suhao Yu, Yihang Sun, Zhanyang Jin, Jiaying Ye, Dianqi Li, Prasanna Sattigeri, Kamal Youcef-Toumi, Hao Peng
arXiv:2605. 07804v3 Announce Type: replace-cross Abstract: On-policy distillation (OPD) leverages dense teacher rewards to enhance reasoning models.
By Zhicheng Yang, Zhijiang Guo, Yifan Song, Minrui Xu, Yongxin Wang, Yiwei Wang, Xiaodan Liang, Jing Tang