arXiv:2604.27251v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicite...
By Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Yuxiang Zhou, Maria Liakata, Nikolaos Aletras
The paper introduces Many-Tier Instruction Hierarchy (ManyIH), a new framework for resolving conflicts among instructions with arbitrarily many privilege levels in large language model agents. It presents ManyIH-Bench, a benchmark featuring 853 agentic tasks that require navigating up to 12 levels of conflicting instructions across 46 real-world agents. Experiments show current models achieve only about 40% accuracy when instruction conflict scales, highlighting a gap in fine-grained, scalable conflict resolution.
By Jingyu Zhang, Tianjian Li, William Jurayj, Hongyuan Zhan, Benjamin Van Durme, Daniel Khashabi
arXiv:2608. 03291v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process.
By Shashwat Sourav, Aishwarya Balwani
arXiv:2607. 04562v1 Announce Type: new Abstract: Large language models (LLMs) generate fluent outputs that can be wrong.
By MY Pitsane, Hope Mogale
arXiv:2607. 16345v1 Announce Type: cross Abstract: Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task.
By Tejas Singh Anand, Yuet Ying Christina Wang, Wanting Jiang, Steve Masson, Tian Zheng, Bingjie Zhou
The paper introduces CONFLICTGUI, a benchmark that tests GUI agents on instruction-internal and instruction‑GUI context conflicts, revealing that many agents over‑comply with infeasible instructions. To address this, the authors propose CONFLICTGUARD, an inference‑time framework that couples a feasibility verification protocol with a conditional action modulation mechanism, enabling agents to assess instruction logic and GUI evidence before acting. Experiments on five agents show that CONFLICTGUARD significantly improves conflict‑task success while maintaining normal task performance.
By Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng, Zheng Wu, Yansi Li, Chuanbiao Song, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.
By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran
arXiv:2511. 04694v5 Announce Type: replace-cross Abstract: As large language model (LLM) based systems take on high-stakes roles in real-world decision-making, they must reconcile competing instructions from multiple sources within a single prompt context.
By Zishuo Zheng, Vidhisha Balachandran, Chan Young Park, Faeze Brahman, Sachin Kumar
arXiv:2606. 05976v2 Announce Type: replace Abstract: Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from external sources.
By Kuan-Yen Chen, Fang-Yi Su, Shih-Yen Lin, Bao Li, Jung-Hsien Chiang
The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.
By Esmail Gumaan
The paper introduces ELCD, a latent conflict detector that verifies LLM outputs after generation to catch instruction conflicts that static input checks miss. ELCD builds a hidden-state representation from the final-token embedding and the mean-pooled response embedding, then trains a pairwise margin ranking objective to distinguish compliant from drifting responses. Experiments on five large language models show ELCD outperforms baselines, boosting PR-AUC for Llama‑2‑7B by ~30 percentage points and cutting FPR95 for Mistral‑7B to 2.67%.
By Mingyu Ma, Yuxin Wu, Jingbo Wang, Tianxiao Huang, Leixin Sun, Xiaochuan Shi
arXiv:2609.01600v1 Announce Type: cross
Abstract: Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a lo...
By Damien Sileo, Dimitri Kachler