arXiv AI

Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces

arXiv:2604. 13780v2 Announce Type: replace-cross Abstract: Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a penalty on the divergence from a reference policy.

arXiv AI
Aug 18

Q-Regularized Generative Auto-Bidding: From Suboptimal Trajectories to Optimal Policies

arXiv:2601. 02754v3 Announce Type: replace-cross Abstract: With the rapid development of e-commerce, auto-bidding has become a key asset in optimizing advertising performance under diverse advertiser environments.

By Mingming Zhang, Na Li, Zhuang Feiqing, Hongyang Zheng, Jiangbing Zhou, Wang Wuyin, Sheng-jie Sun, XiaoWei Chen, Junxiong Zhu, Lixin Zou, Chenliang Li
arXiv AI
Jun 18

Pareto Q-Learning with Reward Machines

arXiv:2606. 19134v1 Announce Type: cross Abstract: We present Pareto Q-Learning with Reward Machines (PQLRM), a multi-objective reinforcement learning algorithm for tasks whose reward structure is specified by a set of reward machines (RMs).

By Arnaud Lequen, Cl\'ement Legrand-Lixon, L\'eo Sauli\`eres
arXiv AI
Sep 11

Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

The paper proposes a new technique that averages the logits of a frozen reference policy (such as one obtained via supervised fine‑tuning) and a trainable policy, integrating this approach into Group Relative Policy Optimization (GRPO). Unlike methods that use Kullback–Leibler regularization or a critic, the proposed method couples the policies solely through logit averaging, aiming to combine the reasoning strengths of the trainable policy with the formatting benefits of the reference policy. Experiments on MATH, cn‑k12, and MMLU demonstrate that this approach achieves higher or comparable accuracy to the standard KL‑regularized GRPO.

By Xingwei Gan, Ying Zhu
arXiv AI
Jul 7

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy

arXiv:2607. 03126v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult.

By Zijun Xie, Yuyang You, Yongzhi Li, Enlei Gong, Zeyu Chen, Quan Chen, Yanhua Cheng, Peng Jiang, Yadong Mu