arXiv AI

FORTIS: Benchmarking Over-Privilege in Agent Skills

arXiv:2605. 09163v3 Announce Type: replace Abstract: Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution.

arXiv AI
Sep 10

AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing

The paper introduces AgentLeak, a black‑box attack that clones the task‑solving capabilities of a strong LLM agent onto a weaker one by exploiting differences between successful and failed executions. Unlike prior skill‑stealing methods that only recover explicit skill artifacts, AgentLeak identifies and incorporates missing procedural behaviors, boosting task pass rates by over 40% and closing more than 80% of the capability gap across 20 scenarios. The study demonstrates that observable execution behavior can leak proprietary procedural knowledge, posing a confidentiality risk for LLM agents.

By Xiaoting Lyu, Yuhong Wu, Yufei Han, Shichang Liu, Liang Zhang, Bin Wang, Bin Wang, Xiaobo Ma, Wei Wang
arXiv AI
Sep 25

Progressive Skill Discovery as Access Control for Tool-Using LLM Agents: Structural Governance through Role-Scoped Capability Delivery

The paper introduces skilder, a framework that organizes LLM agent capabilities into role‑scoped bundles of skills, tools, and instructions, with explicit limits. Agents start with a minimal role catalog, discover the roles needed for a task, and receive the associated tools only through a single MCP server, ensuring deterministic enforcement of scope. Experiments on 13 tasks with six models show that skilder’s authorization layer prevents unauthorized tool calls and parameter violations while maintaining flexibility through dynamic cross‑role capability acquisition.

By Michael Stettler, Benjamin Girardet, Jonas Canton, Nicolas Corod
arXiv AI
Sep 10

Many-Tier Instruction Hierarchy in LLM Agents

The paper introduces Many-Tier Instruction Hierarchy (ManyIH), a new framework for resolving conflicts among instructions with arbitrarily many privilege levels in large language model agents. It presents ManyIH-Bench, a benchmark featuring 853 agentic tasks that require navigating up to 12 levels of conflicting instructions across 46 real-world agents. Experiments show current models achieve only about 40% accuracy when instruction conflict scales, highlighting a gap in fine-grained, scalable conflict resolution.

By Jingyu Zhang, Tianjian Li, William Jurayj, Hongyuan Zhan, Benjamin Van Durme, Daniel Khashabi
arXiv AI
5d ago

EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses

EngramBench is a new benchmark designed to evaluate skill evolution in autonomous agents by focusing on genuine capability abstraction rather than solution copying. It includes 30 learning tasks and 13 unseen transfer tasks that require agents to manage complex, multi-hour development cycles with LLM‑simulated users. The study shows that while static skill banks cannot eliminate the need for precise code implementation, they effectively reduce redundant context and cut overall coding time by more than 55%.

By Zhixuan Tan, Pengjie Gu, Zhao Li, Yihan Hu, Xu He, Dong Li, Jianye Hao
arXiv AI
Aug 20

Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents

The paper introduces a post‑training framework that teaches a 4B‑parameter language model to exercise task‑conditioned authority in executable terminal and Model Context Protocol (MCP) environments. By auditing each action across six risk dimensions with deterministic verifiers and optimizing for task‑specific excess‑privilege values, the authors achieve 98.48% safe success and reduce excess‑authority errors from 4.56% to 0.79% on held‑out tasks. The study also demonstrates capability retention, prompt‑directed improvement, and generalization over a 400‑task continuation test.

By Alexander Tu, Michael Tu
arXiv AI
Aug 19

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It evaluates candidate skills against source evidence and nine safety properties, then tests them in a controlled environment to capture execution traces and identify failures. The system iteratively refines skills based on these results, achieving high precision in vulnerability detection and significantly improving task performance and security rates.

By Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang