arXiv AI By Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi, Sunishchal Dev, Callum Stuart McDougall, Anusha Mujumdar

How Language Models Choose Sides: Internal Representations of Instruction Hierarchy

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
2d ago

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

The paper investigates how large language models balance instruction-following with pattern completion when the two objectives conflict. By creating dialogues where a user instruction to act in a target way T is opposed by assistant turns that demonstrate a competing pattern P, the authors measure instruction-following rates across 13 models and 16 instructions over up to 50 turns. Results show wide variability (1%–99%) in instruction adherence, with robustness influenced by instruction content, output format, and chain-of-thought reasoning, but overall instruction-following remains brittle under induction pressure.

By Carolina Camassa, Derek Shiller