arXiv AI By Yecheng Wu, Song Han, Han Cai

Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition

Read the original on arXiv AI →

Lightning Weave is a post‑training framework that composes independently learned accuracy and efficiency capabilities of reasoning models into a single student model via on‑policy distillation. It extracts policy shifts from pre‑trained models, aligns log‑ratio shifts at shared token states, and uses Tilted‑Target DOPD to create stable learning targets, allowing anchor pairs to score cached trajectories once without running multiple live models. Experiments on mathematics and code benchmarks show significant gains, such as raising Qwen3.5‑4B’s HMMT 2025 accuracy from 59.2% to 64.0% while reducing response tokens by 10.7%, and improving LiveCodeBench v5 accuracy from 41.7% to 54.2% with 9.6% fewer tokens.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

arXiv:2508. 09883v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving.

By Xiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai, Dengfeng Liu, Qiaobo Hao, Huang Liu, Zhiguo Yang, Ji Xie, Ninglun Gu, Jin Yang, Kailai Zhang, Yelun Bao, Jun Wang
arXiv AI
Jun 16

VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models

arXiv:2606. 16140v1 Announce Type: new Abstract: This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly small-model regime.

By Sen Xu, Shixi Liu, Wei Wang, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Xin Zhou, Junlin Zhang
arXiv Machine Learning
Oct 2

Activation-Conditioned Self-Distillation

Activation-Conditioned Self-Distillation (ACSD) is a new on‑policy self‑distillation method that uses a frozen copy of the base model to extract a steering vector by contrasting activations from self‑generated trajectories that reach verified correct answers with all other trajectories. The student learns from next‑token distributions on its own prefixes, without needing reference text or teacher parameter updates, and is used alone at inference. Across five models, ACSD achieves the highest mean accuracy on four mathematical benchmarks, with notable gains on DeepSeek‑R1‑0528‑Qwen3‑8B and LiveCodeBench v6 compared to the OPSD baseline.

By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv Computation and Language
Sep 28

Recursive Self-Improvement via On-Policy Distillation for Reasoning

The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.

By Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz