arXiv AI By Dhriti Krishnan, Tejas Goyal, Jaromir Savelka

Towards Spec Learning: Inference-Time Alignment from Preference Pairs

Read the original on arXiv AI →

arXiv:2606. 24004v1 Announce Type: cross Abstract: Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful inspection of the model's responses.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 28

Instruction Quality Matters: Refining Instructions for Effective Preference Learning

The paper investigates how the quality of instructions used to generate response pairs affects preference learning for language models. It shows that low‑quality or ambiguous instructions limit the range of response quality, weakening preference signals, and introduces an instruction‑refinement pipeline that improves data quality without discarding examples. Experiments across models and benchmarks demonstrate that refining instructions leads to better alignment and complements other data‑improvement methods.

By Seohyeong Lee, Hwaran Lee, Buru Chang
arXiv AI
Jun 16

SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data

arXiv:2606. 16276v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no longer governed by a single universal notion of safety or helpfulness, but instead by provider- or application-specific model specifications.

By Wenjie Wang, Yue Huang, Zhengqing Yuan, Han Bao, Shiyi Du, Yuchen Ma, Yue Zhao, Yanfang Ye, Xiangliang Zhang