arXiv Computation and Language By Seohyeong Lee, Hwaran Lee, Buru Chang

Instruction Quality Matters: Refining Instructions for Effective Preference Learning

Read the original on arXiv Computation and Language →

The paper investigates how the quality of instructions used to generate response pairs affects preference learning for language models. It shows that low‑quality or ambiguous instructions limit the range of response quality, weakening preference signals, and introduces an instruction‑refinement pipeline that improves data quality without discarding examples. Experiments across models and benchmarks demonstrate that refining instructions leads to better alignment and complements other data‑improvement methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.