When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2608.30902v1 Announce Type: new Abstract: Adapting large language models to user-specific preferences is often constrained by the cost of human annotation, making preference optimisation imprac...
arXiv:2509. 22851v4 Announce Type: replace-cross Abstract: Margin-based optimization is fundamental to improving generalization and robustness in classification tasks.
arXiv:2605.16339v2 Announce Type: replace Abstract: Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit prefer...
arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
arXiv:2503. 00539v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs).
arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.