arXiv Machine Learning By Junlong Jia, Jiang Zhou, Ziyang Chen, Xing Wu, Chaochen Gao, TingHao Yu, Feng Zhang, Songlin Hu

PolicyLong: Towards On-Policy Context Extension

Read the original on arXiv Machine Learning →

PolicyLong introduces a dynamic on‑policy approach to constructing long‑context data for large language models, addressing the off‑policy gap of previous single‑pass methods. By repeatedly re‑screening data using the model’s current entropy landscape, it creates a self‑curriculum that aligns training distribution with evolving model capabilities. Experiments on RULER, HELMET, and LongBench‑v2 demonstrate consistent performance gains, especially at longer contexts, outperforming EntropyLong and NExtLong.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 11

Adaptive Supervised Anchoring for On-Policy Self-Distillation

arXiv:2608. 07935v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student.

By Meilin Yang (Renmin University of China, Beijing, China), Zixuan Ding (Renmin University of China, Beijing, China), Jianhao Nie (Renmin University of China, Beijing, China), Weite Zhang (Renmin University of China, Beijing, China), Yuxin Zhang (Renmin University of China, Beijing, China), Zhiming Shao (Renmin University of China, Beijing, China), Li Yu (Renmin University of China, Beijing, China), Zhe Fu (Renmin University of China, Beijing, China)
arXiv Computation and Language
4d ago

When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

The paper introduces a holistic framework for large language model self‑evolution that uses learnable information gain to assess the novelty of each training round. Information gain is theoretically linked to the Kullback‑Leibler divergence and entropy change between successive data distributions, and practically estimated by fitting a small language model and scoring new data with negative log‑likelihood. The proposed ATRI method reweights samples within a round and stops training across rounds when information gain is low, and experiments on popular datasets show its effectiveness.

By Chenxu Wang, Chaozhuo Li, Xinze Shi, Songyang Liu, Kyrie You Wu, Ziluowen Luo, Shun Zhang, Chenxi Li, Litian Zhang