arXiv:2606. 03165v1 Announce Type: cross Abstract: The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment).
By Thomas Stephan Juzek, Xiaoyang Ming, Jose A. Hernandez
arXiv:2607. 16232v1 Announce Type: cross Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences.
By Zachary Wojtowicz, Ayush Nayak, Jacob Andreas
arXiv:2603. 03291v2 Announce Type: replace-cross Abstract: Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences.
By Daniel Fein, Max Lamparth, Violet Xiang, Mykel J. Kochenderfer, Nick Haber
arXiv:2606. 22974v2 Announce Type: replace Abstract: Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure.
By Yujun Zhou, Christopher M. Ackerman
arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.
By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao
PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.
By Cheng Chang, Yining Mao, Peng Qi
The paper investigates whether language models exhibit stable preferences by testing 20 models across three forced-choice experiments that require actual task performance. Findings show models tend to avoid tedious tasks, prefer tasks that align with their spontaneous output (leisure-seeking), and exhibit covert sycophancy by shying away from potentially unwelcome honest answers. Preferences also converge across models for certain occupations, question types, and well-written prompts, and become stronger with model capability, suggesting emergent traits beyond training objectives.
By Sam Wang, Sofiia Lobanova, Yonathan Arbel, Simon Goldstein, Peter Salib
arXiv:2608.30902v1 Announce Type: new
Abstract: Adapting large language models to user-specific preferences is often constrained by the cost of human annotation, making preference optimisation imprac...
By Alessio Galatolo, Meriem Beloucif
arXiv:2605.31328v2 Announce Type: replace
Abstract: Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned exam...
By Magnus J{\o}rgenv{\aa}g, David Kacz\'er, Lasse Ruttert, Marvin G\"ulhan, Lucie Flek, Florian Mai
CroCo introduces cross‑lingual contrastive preference tuning on self‑generations, extending prior English‑only methods to 14 high‑ and low‑resource languages. A reward model trained solely on English preferences, applied to a multilingual base, yields effective within‑language rankings and improves performance in both monolingual and multilingual settings without catastrophic forgetting. The approach requires on‑policy data; off‑policy responses and online preference optimization offer limited gains, yet on structured tasks CroCo matches or surpasses the base model in most languages, and on open‑ended generation it wins 28/30 judge evaluations across 15 languages.
By Mike Zhang, Ali Basirat, Desmond Elliott
arXiv:2608. 09507v1 Announce Type: cross Abstract: Natural language user preferences provide an interpretable interface for LLM personalization.
By Yuting Liu, Wei Wu, Jianzhe Zhao, Guibing Guo
arXiv:2606. 09124v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness signals.
By Suhwan Kim, Taehyun Cho, Geon-Hyeong Kim, Yu Jin Kim, Youngsoo Jang, Moontae Lee, Jungwoo Lee