arXiv:2606. 01382v1 Announce Type: cross Abstract: Preference alignment is central to improving large language models, but standard reward-based formulations can be restrictive when human preferences are cyclic, non-transitive, or otherwise not representable by a scalar reward.
By Tianlong Nan, Xiaopeng Li, Christian Kroer, Tianyi Lin
The paper investigates preference-based bandits where a learner selects pairs of arms and receives binary preference feedback modeled by Bradley–Terry. It introduces the locally sensitive eluder dimension, a new complexity measure for logistic preference feedback, and proposes the GINOP algorithm that uses log-loss confidence sets to balance optimism and exploration. The authors prove a first-order regret bound showing that learning with preference feedback can be as statistically efficient as learning from direct rewards, and they validate their theory with empirical experiments.
By Ahmed Ben Yahmed (CREST, ENSAE Paris, FAIRPLAY), Marc Abeille (FAIRPLAY), Cl\'ement Calauz\`enes (FAIRPLAY)
arXiv:2607. 26358v1 Announce Type: new Abstract: Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy.
By Keegan Harris, Brian W. Lee, Ian Waudby-Smith, Philip Amortila, Nika Haghtalab, Michael I. Jordan
Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient.
The paper presents a new scaling law for reward optimization in AI alignment, showing that performance scales as Θ(√min{log(M), K}), where M is the number of preference comparisons used to train a proxy reward model and K is the KL‑divergence budget relative to a reference policy. The authors derive this law using an information‑theoretic model, prove its tightness, and validate it with extensive experiments involving a 70B gold reward model and smaller proxy models (0.6B–4B). The empirical results demonstrate a strong fit (R² 97–99 %) across different model sizes, noise levels, and optimization methods, suggesting that reward optimization behaves like a simple selection task over IID Gaussian variables with noisy feedback.
By Ali Aouad, Aymane El Gadarri, Vivek F. Farias
arXiv:2607. 26094v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models.
By Yunpeng Chu
arXiv:2606. 18111v1 Announce Type: cross Abstract: Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives.
By Umer Siddique, Peilang Li, Yongcan Cao
arXiv:2609.37209v1 Announce Type: new
Abstract: Pairwise losses are increasingly used for reward learning even when pointwise rewards are observed, with mixed empirical results. When and why do pairw...
By Junghyun Lee, Minsoo Ha, Sanghwa Kim, Yeongjong Kim, Eunjee Lee, Seiyun Shin, Kwang-Sung Jun
arXiv:2403.05006v2 Announce Type: replace-cross
Abstract: Pluralistic alignment requires learning from feedback that reflects persistent and potentially conflicting stakeholder preferences while ulti...
By Huiying Zhong, Tianwei Gao, Zhiwei Steven Wu, Linjun Zhang, Weijie J. Su, Zhun Deng
arXiv:2501.06926v5 Announce Type: replace
Abstract: Double reinforcement learning (DRL) provides efficient off-policy inference for policy values in nonparametric Markov decision processes (MDPs), bu...
By Lars van der Laan, David Hubbard, Allen Tran, Nathan Kallus, Aur\'{e}lien Bibaut
arXiv:2606. 19117v1 Announce Type: cross Abstract: Offline policy learning has received growing attention in causal inference.
By Yiyan Huang, Cheuk Hang Leung, Qi Wu, Zhiheng Zhang
arXiv:2607. 11432v1 Announce Type: new Abstract: In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert.
By Simone Drago, Marco Mussi, Leonardo Bianconi, Alberto Maria Metelli