Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient.
arXiv:2602. 18037v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) or Verifiable Rewards (RLVR) are two key steps in the post-training of modern Language Models (LMs).
By Johannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi Sugiyama
arXiv:2606. 01382v1 Announce Type: cross Abstract: Preference alignment is central to improving large language models, but standard reward-based formulations can be restrictive when human preferences are cyclic, non-transitive, or otherwise not representable by a scalar reward.
By Tianlong Nan, Xiaopeng Li, Christian Kroer, Tianyi Lin
The paper presents a new scaling law for reward optimization in AI alignment, showing that performance scales as Θ(√min{log(M), K}), where M is the number of preference comparisons used to train a proxy reward model and K is the KL‑divergence budget relative to a reference policy. The authors derive this law using an information‑theoretic model, prove its tightness, and validate it with extensive experiments involving a 70B gold reward model and smaller proxy models (0.6B–4B). The empirical results demonstrate a strong fit (R² 97–99 %) across different model sizes, noise levels, and optimization methods, suggesting that reward optimization behaves like a simple selection task over IID Gaussian variables with noisy feedback.
By Ali Aouad, Aymane El Gadarri, Vivek F. Farias
arXiv:2608. 01743v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model.
By Li Wang, Xiaodong Lu, Xiaohan Wang, Jiajun Chai, Wei Lin, Tianhao Peng, Guojun Yin
arXiv:2609.36945v1 Announce Type: new
Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version...
By Zhiwei Wang, Yanxi Chen, Yaliang Li, Bolin Ding
arXiv:2606. 03962v1 Announce Type: cross Abstract: Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward.
By Anthony GX-Chen, Ankit Anand, Gheorghe Comanici, Zaheer Abbas, Eser Ayg\"un, David Smalling, Shibl Mourad, Doina Precup, Andr\'e Barreto, Mark Rowland
arXiv:2511. 23310v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective paradigm for post-training large language models, yet the design of its baselines and learning-rate schedules remains largely heuristic.
By Zixun Huang, Jiayi Sheng, Zeyu Zheng
The paper investigates how different estimators of the reverse Kullback–Leibler (KL) divergence used as a regularization term in reinforcement learning (RL) training of large language models (LLMs) affect training stability and downstream performance. By analyzing gradient bias across various estimator configurations, the authors demonstrate that biased gradients can cause training instabilities, while unbiased configurations improve performance on both in‑domain and out‑of‑domain tasks. Experiments on Qwen2.5‑7B, Llama‑3.1‑8B‑Instruct, and Qwen3‑4B‑Instruct‑2507 confirm these findings and show that KL regularization also stabilizes off‑policy RL training in asynchronous setups.
By Vedant Shah, Johan Obando-Ceron, Vineet Jain, Brian Bartoldson, Bhavya Kailkhura, Sarthak Mittal, Glen Berseth, Pablo Samuel Castro, Yoshua Bengio, Esmeralda S. Whitammer, Moksh Jain, Siddarth Venkatraman, Aaron Courville
arXiv:2607. 28582v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort.
By Jiawei Xu, Minghui Liu, Juzheng Zhang, Tom Goldstein, Furong Huang
arXiv:2608. 09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization.
By Ting Zhou, Zhenqing Ling, Daoyuan Chen, Qianli Shen, Yilun Huang, Ying Shen, Yaliang Li
arXiv:2609.36393v1 Announce Type: cross
Abstract: Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit...
By Muhang Tian, Sherry Yang