arXiv AI

Towards Spec Learning: Inference-Time Alignment from Preference Pairs

arXiv:2606. 24004v1 Announce Type: cross Abstract: Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful inspection of the model's responses.

arXiv Computation and Language
Aug 28

Instruction Quality Matters: Refining Instructions for Effective Preference Learning

The paper investigates how the quality of instructions used to generate response pairs affects preference learning for language models. It shows that low‑quality or ambiguous instructions limit the range of response quality, weakening preference signals, and introduces an instruction‑refinement pipeline that improves data quality without discarding examples. Experiments across models and benchmarks demonstrate that refining instructions leads to better alignment and complements other data‑improvement methods.

By Seohyeong Lee, Hwaran Lee, Buru Chang
arXiv AI
Jun 16

SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data

arXiv:2606. 16276v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no longer governed by a single universal notion of safety or helpfulness, but instead by provider- or application-specific model specifications.

By Wenjie Wang, Yue Huang, Zhengqing Yuan, Han Bao, Shiyi Du, Yuchen Ma, Yue Zhao, Yanfang Ye, Xiangliang Zhang
arXiv AI
Sep 10

DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment

The paper introduces DSPA, a dynamic sparse autoencoder (SAE) steering technique that aligns language model outputs with user preferences during inference, avoiding costly weight updates. DSPA constructs a conditional-difference map from preference triples to adjust token-active latents, improving MT‑Bench scores and matching AlpacaEval performance on models like Gemma‑2 and Qwen3 while preserving accuracy. It demonstrates robustness with limited preference data, outperforms the two‑stage RAHF‑SCIT pipeline in FLOPs, and reveals that preference directions are largely driven by discourse and stylistic cues.

By James Wedgwood, Aashiq Muhamed, Mona T. Diab, Virginia Smith
arXiv AI
2d ago

Gradient-Aligned Pair Selection for Personalized Preference Optimization

The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.

By Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou