arXiv AI By Xu Wan, Wenyue Xu, Shengjie Zhao, Mingyang Sun

Mitigating the Length-Scaling Tax with Online Distillation

Read the original on arXiv AI →

The paper introduces Length Self-Distillation (LSD) to address the length‑scaling tax (LST) that occurs during reinforcement‑learning post‑training, where models produce unnecessarily verbose responses to already‑solved prompts. LSD routes solved prompts to an on‑policy distillation process while keeping the original RL objective for unsolved prompts, using an exponential moving average of the online policy as its teacher. Experiments show LSD matches or surpasses RL performance while reducing LST from 19.0% to –3.7% on single‑turn reasoning and from 31.4% to 13.7% on multi‑turn agentic tasks, thereby maintaining concise responses on easy queries while still enabling exploration on difficult ones.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

OPDSearch+ introduces a two‑stage distillation framework for search‑augmented reasoning that eliminates the need for task‑specific teacher fine‑tuning. In the first stage, a frozen off‑the‑shelf instruct model guides a student through live search interactions using a per‑position forward KL objective, transferring reasoning decomposition and evidence integration skills. The second stage refines this student with reinforcement learning, achieving performance surpassing RL alone and outperforming all prior 3B‑parameter baselines on seven QA benchmarks, including 13.1% improvement on HotpotQA and 8.5% on 2WikiMultihopQA.

By Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng, Yuhang Mu, Wenchao Du, Yiming Wang