arXiv:2608.30902v1 Announce Type: new
Abstract: Adapting large language models to user-specific preferences is often constrained by the cost of human annotation, making preference optimisation imprac...
By Alessio Galatolo, Meriem Beloucif
arXiv:2509. 22851v4 Announce Type: replace-cross Abstract: Margin-based optimization is fundamental to improving generalization and robustness in classification tasks.
By Yaswanth Chittepu, Prasann Singhal, Greg Durrett, Scott Niekum
arXiv:2605.16339v2 Announce Type: replace
Abstract: Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit prefer...
By Shunchang Liu, Xin Chen, Belen Martin Urcelay, Francesco Croce
arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
By Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf
arXiv:2503. 00539v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs).
By Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic
arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.
By Hyung Gyu Rho
The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.
By Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
Data-DPO is a target model‑oriented supervised fine‑tuning data selection method that uses one‑step probing of the target model to generate pairwise data preferences, trains a lightweight reward model to capture these preferences, and then selects a training subset by combining target‑model preference, external quality scores, and marginal diversity. Experiments on Vision‑Flan and LLaVA‑CoT demonstrate that Data‑DPO consistently outperforms existing data selection baselines across multiple data budgets and even surpasses full data training performance.
By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu
One-step text-to-image generators are attractive for deployment because they generate an image with a single forward pass, but preference finetuning them remains difficult: standard alignment methods often rely on policy likelihoods, denoising trajectories, differentiable reward gradients, or test-time optimization. We propose Drifting Preference Optimization (DrPO), an online preference-finetuning method for deterministic one-step generators.
The paper introduces TESS, a scalable data‑selection framework that replaces per‑sample weights with a selection network to improve transferability across datasets and model sizes. It identifies instability in existing meta‑learning for training‑data selection (MTS) due to weight suppression and overreliance on easy features, and proposes a Pointwise Value Matching objective to address these issues. Experiments on large language model safety and instruction tuning show strong transfer from subsets to full corpora and from smaller to larger models.
By Zilin Du, Bowen Yang, Boyang Albert Li
The paper introduces Geometric Anchor Preference Optimization (GAPO), a method that replaces the static reference policy in Direct Preference Optimization with a dynamic, geometry-aware anchor—a small adversarial perturbation of the current policy. GAPO uses this anchor to adaptively reweight preference pairs based on local sensitivity, and defines an Anchor Gap that approximates worst‑case local margin degradation. Experiments show that GAPO improves robustness to noisy supervision while matching or surpassing existing LLM alignment and reasoning benchmarks.
By Youngjae Cho, Jongsuk Kim, Ji-Hoon Kim
arXiv:2609.30840v1 Announce Type: cross
Abstract: One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit ge...
By Austin Wang, Ziheng Cheng, Lexing Ying