Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2509. 23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences.
Robust Nash Alignment introduces a game-theoretic framework that seeks a policy with a high worst-case win rate against both an adversarial competitor and any preference kernel within an ambiguity set around a nominal preference. The authors propose a four-player primal-dual proxy game and an optimistic mirror descent-ascent algorithm to efficiently optimize this robust objective, proving convergence guarantees and demonstrating improved performance in controlled tabular games and LLM alignment experiments.
arXiv:2503.00030v3 Announce Type: replace-cross Abstract: Self-play-based policy optimization has emerged as an effective approach for fine-tuning large language models (LLMs), formulating preference...
arXiv:2606. 01561v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO).
arXiv:2608. 15402v1 Announce Type: new Abstract: Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation.
Swiss-Knife is a framework that extends decode‑time alignment for frozen language models by treating the alignment specification as a runtime object. It introduces hot‑swappable scoring blades, a batch normaliser, a pairwise aggregation operator, and a selection rule, and characterises admissible aggregation operators with a representation theorem. In experiments, Swiss‑Knife paired with DPO‑LoRA blades and an uncertainty‑aware pairwise tournament outperforms six existing decode‑time methods, achieving a higher harmonic F1 score, lower refusal rate, and faster objective reconfiguration.