OpenAI Blog

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Read the original on OpenAI Blog →

Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

arXiv AI
Sep 10

Many-Tier Instruction Hierarchy in LLM Agents

The paper introduces Many-Tier Instruction Hierarchy (ManyIH), a new framework for resolving conflicts among instructions with arbitrarily many privilege levels in large language model agents. It presents ManyIH-Bench, a benchmark featuring 853 agentic tasks that require navigating up to 12 levels of conflicting instructions across 46 real-world agents. Experiments show current models achieve only about 40% accuracy when instruction conflict scales, highlighting a gap in fine-grained, scalable conflict resolution.

By Jingyu Zhang, Tianjian Li, William Jurayj, Hongyuan Zhan, Benjamin Van Durme, Daniel Khashabi