How Language Models Choose Sides: Internal Representations of Instruction Hierarchy
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper investigates how large language models balance instruction-following with pattern completion when the two objectives conflict. By creating dialogues where a user instruction to act in a target way T is opposed by assistant turns that demonstrate a competing pattern P, the authors measure instruction-following rates across 13 models and 16 instructions over up to 50 turns. Results show wide variability (1%–99%) in instruction adherence, with robustness influenced by instruction content, output format, and chain-of-thought reasoning, but overall instruction-following remains brittle under induction pressure.
arXiv:2606. 15420v1 Announce Type: cross Abstract: A constitution tells a language model what to value, but little tells us whether it does.
arXiv:2606. 07808v1 Announce Type: new Abstract: Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the model should obey the highest-privilege applicable instruction.
arXiv:2608. 13921v1 Announce Type: new Abstract: LLM agents increasingly maintain personal memory across sessions, but it can conflict.
arXiv:2607. 24765v1 Announce Type: cross Abstract: Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own prior answer returns as context.
LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks context, time, or s...