Cross-Relational Preference Learning for Better LLM Instruction Following
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper investigates how the quality of instructions used to generate response pairs affects preference learning for language models. It shows that low‑quality or ambiguous instructions limit the range of response quality, weakening preference signals, and introduces an instruction‑refinement pipeline that improves data quality without discarding examples. Experiments across models and benchmarks demonstrate that refining instructions leads to better alignment and complements other data‑improvement methods.
arXiv:2606. 24004v1 Announce Type: cross Abstract: Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful inspection of the model's responses.
arXiv:2509. 23982v2 Announce Type: replace-cross Abstract: Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences.
arXiv:2609.17019v1 Announce Type: new Abstract: While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference...
arXiv:2607. 22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs).
arXiv:2503. 06573v3 Announce Type: replace-cross Abstract: Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge.