arXiv Machine Learning

PolicyLong: Towards On-Policy Context Extension

PolicyLong introduces a dynamic on‑policy approach to constructing long‑context data for large language models, addressing the off‑policy gap of previous single‑pass methods. By repeatedly re‑screening data using the model’s current entropy landscape, it creates a self‑curriculum that aligns training distribution with evolving model capabilities. Experiments on RULER, HELMET, and LongBench‑v2 demonstrate consistent performance gains, especially at longer contexts, outperforming EntropyLong and NExtLong.

arXiv Machine Learning
Aug 11

Adaptive Supervised Anchoring for On-Policy Self-Distillation

arXiv:2608. 07935v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student.

By Meilin Yang (Renmin University of China, Beijing, China), Zixuan Ding (Renmin University of China, Beijing, China), Jianhao Nie (Renmin University of China, Beijing, China), Weite Zhang (Renmin University of China, Beijing, China), Yuxin Zhang (Renmin University of China, Beijing, China), Zhiming Shao (Renmin University of China, Beijing, China), Li Yu (Renmin University of China, Beijing, China), Zhe Fu (Renmin University of China, Beijing, China)
arXiv Computation and Language
4d ago

When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

The paper introduces a holistic framework for large language model self‑evolution that uses learnable information gain to assess the novelty of each training round. Information gain is theoretically linked to the Kullback‑Leibler divergence and entropy change between successive data distributions, and practically estimated by fitting a small language model and scoring new data with negative log‑likelihood. The proposed ATRI method reweights samples within a round and stops training across rounds when information gain is low, and experiments on popular datasets show its effectiveness.

By Chenxu Wang, Chaozhuo Li, Xinze Shi, Songyang Liu, Kyrie You Wu, Ziluowen Luo, Shun Zhang, Chenxi Li, Litian Zhang
arXiv Machine Learning
Jun 5

Extreme Region Policy Distillation

arXiv:2605. 25582v2 Announce Type: replace Abstract: Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories after a single update, while off-policy reuse introduces distribution mismatch that existing trust-region techniques mitigate primarily by enforcing conservative optimization, often leaving rich training signals underutilized.

By Changyu Chen, Xiting Wang, Rui Yan
arXiv Machine Learning
Sep 14

MInTRL: Off-policy Intervention can boost On-policy RL

MInTRL (Minimal Intervention Reinforcement Learning) expands exploration in on-policy reinforcement learning by inserting sparse, local corrections into rollouts via a judge-intervention policy. These interventions replace erroneous suffixes and immediately return control to the main policy, allowing the agent to explore beyond its natural trajectory while maintaining on-policy data. The method uses a sequence-level advantage-regression objective, avoiding importance sampling, and demonstrates superior performance on math and code benchmarks compared to standard on-policy and off-policy baselines.

By Mingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong