arXiv AI By Tokio Kajitsuka, Ukyo Honda, Sho Takase

Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Sep 2

When Compression Helps and When It Hurts: Condition-Aware Analysis of Chain-of-Thought Distillation

The paper studies how to compress Chain-of-Thought (CoT) reasoning traces for smaller models. It examines three compression dimensions—importance criterion, restructuring level, and compression budget—across Math and General domains and Long/Short CoT regimes. Findings show that step-level pruning works best for shared reasoning backbones, token-level pruning needs symbol-aware signals, domain-specific restructuring effects differ, and training-time compression may not reduce inference cost, especially for Long-CoT students.

By Siyang Lyu, Xinghao Chen, Zhijing Sun, Tong Liu, Dawei Zhu, Xiaoyu Shen
arXiv Computation and Language
Sep 18

What Does Privileged Information Add to On-Policy Self-Distillation?

The paper investigates how privileged information—such as a teacher’s full solution or reasoning trace—affects on‑policy self‑distillation (OPSD) in language models. Using the AMPLE‑Math benchmark, the authors compare distillation with and without extra teacher views, finding that reference‑free distillation explains most gains for Qwen3‑1.7B, while additional references provide modest benefits, especially for polished solutions. The study also shows that the impact of privileged data depends on the student’s training regime and that altering token‑level supervision can leave student behavior largely unchanged.

By XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang, Tat-Seng Chua