Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper studies how to compress Chain-of-Thought (CoT) reasoning traces for smaller models. It examines three compression dimensions—importance criterion, restructuring level, and compression budget—across Math and General domains and Long/Short CoT regimes. Findings show that step-level pruning works best for shared reasoning backbones, token-level pruning needs symbol-aware signals, domain-specific restructuring effects differ, and training-time compression may not reduce inference cost, especially for Long-CoT students.
arXiv:2608. 08294v1 Announce Type: new Abstract: Knowledge distillation trains a smaller student to match the outputs of a larger teacher.
arXiv:2608. 11829v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning.
arXiv:2609.36246v1 Announce Type: new Abstract: We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressivel...
The paper investigates how privileged information—such as a teacher’s full solution or reasoning trace—affects on‑policy self‑distillation (OPSD) in language models. Using the AMPLE‑Math benchmark, the authors compare distillation with and without extra teacher views, finding that reference‑free distillation explains most gains for Qwen3‑1.7B, while additional references provide modest benefits, especially for polished solutions. The study also shows that the impact of privileged data depends on the student’s training regime and that altering token‑level supervision can leave student behavior largely unchanged.
arXiv:2607. 26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs.