arXiv AI By Xiaodong Wang, Peixi Peng

Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning

Read the original on arXiv AI →

The paper introduces the Flow Moment, a reasoning pattern marked by sustained, process‑confirming verbalizations, contrasting with the revision‑oriented Aha Moment. It proposes Flow‑CoT, a rewritten version of reasoning traces that preserves content while highlighting Flow Markers, and uses it as auxiliary supervision in on‑policy self‑distillation (OPSD). The authors further present Aha‑Flow Distillation (AFD), a dual‑mode extension of OPSD that pairs concise solution‑based supervision (Aha branch) with rewritten Flow‑CoT under a confident reasoning instruction (Flow branch). Experiments on AIME25 and HMMT25 with Qwen3‑8B and Qwen3‑4B models show consistent performance gains, and controlled ablations confirm that the dual‑mode training structure contributes to the improvement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 18

What Does Privileged Information Add to On-Policy Self-Distillation?

The paper investigates how privileged information—such as a teacher’s full solution or reasoning trace—affects on‑policy self‑distillation (OPSD) in language models. Using the AMPLE‑Math benchmark, the authors compare distillation with and without extra teacher views, finding that reference‑free distillation explains most gains for Qwen3‑1.7B, while additional references provide modest benefits, especially for polished solutions. The study also shows that the impact of privileged data depends on the student’s training regime and that altering token‑level supervision can leave student behavior largely unchanged.

By XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang, Tat-Seng Chua
arXiv AI
Jul 3

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

arXiv:2607. 02234v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-level supervision on the student's own generated trajectories.

By Zhanming Shen, Jintao Tong, Shaotian Yan, Chen Shen, Hao Chen, Wentao Ye, Xiaomeng Hu, Rui Miao, Haobo Wang, Junbo Zhao, Gang Chen, Jieping Ye
arXiv AI
Jul 22

Fluid Reasoning Representations

arXiv:2602. 04843v2 Announce Type: replace Abstract: Frontier large language models increasingly solve complex tasks involving abstract concepts through extended test-time thinking.

By Dmitrii Kharlapenko, Terry Jingchen Zhang, Arth Singh, Alessandro Stolfo, Arthur Conmy, Mrinmaya Sachan, Zhijing Jin