Inference-Time Nash Alignment
arXiv:2609.08082v1 Announce Type: new Abstract: Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to th...
Robust Nash Alignment introduces a game-theoretic framework that seeks a policy with a high worst-case win rate against both an adversarial competitor and any preference kernel within an ambiguity set around a nominal preference. The authors propose a four-player primal-dual proxy game and an optimistic mirror descent-ascent algorithm to efficiently optimize this robust objective, proving convergence guarantees and demonstrating improved performance in controlled tabular games and LLM alignment experiments.
arXiv:2609.08082v1 Announce Type: new Abstract: Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to th...
arXiv:2601.08777v2 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for pe...
arXiv:2607. 01715v1 Announce Type: new Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level.
arXiv:2606. 01382v1 Announce Type: cross Abstract: Preference alignment is central to improving large language models, but standard reward-based formulations can be restrictive when human preferences are cyclic, non-transitive, or otherwise not representable by a scalar reward.
arXiv:2509. 23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences.
arXiv:2505. 10892v2 Announce Type: replace Abstract: Post-training LLMs with RLHF and preference optimization methods (e.
arXiv:2606. 00002v1 Announce Type: new Abstract: Mixed-Integer Linear Programming (MILP) decision engines routinely output nominally optimal plans for high-stakes industrial systems.
arXiv:2512. 15765v3 Announce Type: replace Abstract: Data valuation is a natural framework for understanding which preference datasets matter most when aligning a Large Language Model (LLM) using multiple sources.
The paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method that aligns large language models with human preferences using comparison oracles instead of direct differentiable loss optimization. It provides theoretical convergence guarantees for both offline and online variants under smoothness, gradient sparsity, and oracle compatibility assumptions, and establishes performance bounds under local coverage and in-distribution reward accuracy. Experiments on several LLMs (Mistral, Llama, Gemma-2, Qwen3, Gemma-3) show that ComPO outperforms existing direct alignment methods, achieving higher length-controlled win rates and diagnostics that suggest mitigation of likelihood displacement.
The paper introduces Geometric Anchor Preference Optimization (GAPO), a method that replaces the static reference policy in Direct Preference Optimization with a dynamic, geometry-aware anchor—a small adversarial perturbation of the current policy. GAPO uses this anchor to adaptively reweight preference pairs based on local sensitivity, and defines an Anchor Gap that approximates worst‑case local margin degradation. Experiments show that GAPO improves robustness to noisy supervision while matching or surpassing existing LLM alignment and reasoning benchmarks.
arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.
arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).