arXiv Computation and Language

Combating Instruction Conflict via Energy-Driven Latent Conflict Detection

The paper introduces ELCD, a latent conflict detector that verifies LLM outputs after generation to catch instruction conflicts that static input checks miss. ELCD builds a hidden-state representation from the final-token embedding and the mean-pooled response embedding, then trains a pairwise margin ranking objective to distinguish compliant from drifting responses. Experiments on five large language models show ELCD outperforms baselines, boosting PR-AUC for Llama‑2‑7B by ~30 percentage points and cutting FPR95 for Mistral‑7B to 2.67%.

arXiv AI
Sep 10

Many-Tier Instruction Hierarchy in LLM Agents

The paper introduces Many-Tier Instruction Hierarchy (ManyIH), a new framework for resolving conflicts among instructions with arbitrarily many privilege levels in large language model agents. It presents ManyIH-Bench, a benchmark featuring 853 agentic tasks that require navigating up to 12 levels of conflicting instructions across 46 real-world agents. Experiments show current models achieve only about 40% accuracy when instruction conflict scales, highlighting a gap in fine-grained, scalable conflict resolution.

By Jingyu Zhang, Tianjian Li, William Jurayj, Hongyuan Zhan, Benjamin Van Durme, Daniel Khashabi
arXiv Computation and Language
Sep 18

IHDec: Divergence-Steered Contrastive Decoding for Securing Multi-Turn Instruction Hierarchies

Large Language Models often fail to respect instruction hierarchies in multi-turn settings, sometimes following lower-priority directives over higher ones. The authors formalize this failure using a Jensen‑Shannon Divergence framework and introduce IHDec, a contrastive decoding method that detects hierarchy violations at the token level and suppresses subordinate role influence without any fine‑tuning. Experiments show IHDec outperforms training‑based baselines, maintains overall response quality, improves safety against adversarial prompts, and scales well with larger models.

By Nicole Geumheon Liu, Haeun Jang, Yonghyun Jun, Hwanhee Lee
arXiv AI
Jun 8

SWE-IF: Aligning Code Evaluation with Human Preference

arXiv:2510. 07315v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have catalyzed vibe coding, where users leverage LLMs to generate and iteratively refine code through natural language interactions until it passes their vibe check.

By Ming Zhong, Xiang Zhou, Ting-Yun Chang, Qingze Wang, Nan Xu, Xiance Si, Dan Garrette, Shyam Upadhyay, Jeremiah Liu, Jiawei Han, Benoit Schillings, Jiao Sun
arXiv AI
Aug 20

Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions

The paper introduces a fine-grained method called interactions to analyze prompt sensitivity in large language models (LLMs). By decomposing output scores into nonlinear interactions, the authors show that subtle prompt changes can destabilize these interactions even when overall outputs stay unchanged. They propose an Interaction-based Prompt Sensitivity (IPS) metric and use it to evaluate 50 open-source LLMs, finding that supervised fine‑tuning, larger model scales, dense architectures, and few‑shot learning all reduce prompt sensitivity, primarily by stabilizing low‑order interactions.

By Ruiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei, Wen Shen
Hugging Face Trending Papers
Jul 30

IFHierBench: Hierarchical Instruction Following for Large Language Models

Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the output satisfying specific constraints. Modern deployments increasingly handle the full task in a single LLM call, with one prompt specifying a layered output whose overall artifact, structural sections, and nested fields must each satisfy concrete constraints.

arXiv AI
Jun 4

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

arXiv:2602. 06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks.

By Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee, Matthew Kowal, Nayeema Nonta, Samuel Simko, Stephen Casper, Zhijing Jin, Kellin Pelrine, Sirisha Rambhatla