Autoregressive Direct Preference Optimization
arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.
arXiv:2602. 10286v3 Announce Type: replace Abstract: Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences.
arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.
The paper argues that language models’ intransitive preferences arise from multiple internally consistent latent orderings rather than noise around a single ordering. By demonstrating that a single ordering cannot explain observed inconsistencies and introducing a noise‑augmented mixture Bradley‑Terry model, the authors show that mixtures of orderings better capture preference structure across several models and tasks. A case study on Moral Machine dilemmas further illustrates that models can share latent components even when aggregate preferences differ.
arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.
arXiv:2606. 19744v1 Announce Type: cross Abstract: Aligning language models with human preferences often requires optimising multiple behavioural objectives.
The paper introduces DSPA, a dynamic sparse autoencoder (SAE) steering technique that aligns language model outputs with user preferences during inference, avoiding costly weight updates. DSPA constructs a conditional-difference map from preference triples to adjust token-active latents, improving MT‑Bench scores and matching AlpacaEval performance on models like Gemma‑2 and Qwen3 while preserving accuracy. It demonstrates robustness with limited preference data, outperforms the two‑stage RAHF‑SCIT pipeline in FLOPs, and reveals that preference directions are largely driven by discourse and stylistic cues.
The paper introduces BALIGN, a balanced data selection strategy designed to reduce catastrophic forgetting—referred to as the alignment tax—in large language models during preference-based alignment. By analyzing preference optimization gradients, the authors identify three data-centric features that influence parameter drift: the reference model's log-probability margin, token length differences between chosen and rejected responses, and TF‑IDF similarity to general capability corpora. BALIGN aggregates these features into a composite risk score to filter out high-risk preference samples, thereby preserving foundational capabilities while maintaining alignment gains with minimal computational overhead.
arXiv:2606. 19607v1 Announce Type: new Abstract: Preference-based post-training has become a central paradigm for aligning language models.
arXiv:2605. 12288v3 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences over full sequences even though generation is driven by per-token decisions.
Aligning large language models to human preferences is crucial for real-world deployment but frequently incurs an alignment tax, leading to the catastrophic forgetting of pre-trained general capabilit...
The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.
SCOPE is a framework that calibrates an acceptance threshold for large language models used as pairwise judges, ensuring that the error rate among non-abstained judgments does not exceed a user-specified level α. It introduces Bidirectional Preference Entropy (BPE) to provide a bias-neutral uncertainty signal by querying the judge in both response positions and converting the averaged preference probability into an entropy-based score. Across multiple pairwise judging benchmarks, BPE outperforms standard confidence proxies in calibration and discrimination, while SCOPE consistently meets the target risk bound (empirical FDR ≈0.097–0.099 at α=0.10) and retains substantial coverage, accepting up to 2.4× more judgments under the same risk constraint.
arXiv:2607. 16232v1 Announce Type: cross Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences.