End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where. We present layer-isolated evaluation: a deployed ordering agent is decomposed into a fixed taxonomy of layers (ontology, intent, routing, decomposition, escalation, safety, memory, and cross-cutting envelope/defense), each exercised by its own assertion slice in a deterministic, no-LLM "pure" mode.
arXiv:2609.35889v1 Announce Type: cross
Abstract: Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing...
By Xiaoyu Xu, Zi Liang, Minxin Du, Qipeng Xie, Qingqing Ye, Yuyuan Li, Haibo Hu
The paper examines safety routers—systems that route user requests to different language models—and finds that their performance degrades significantly when evaluated under distribution shift. In standard benchmarks, routers appear effective because the best single model is chosen from the same evaluation data, but when the data distribution changes, the routing advantage diminishes or disappears. The study quantifies this bias across multiple safety corpora, showing that routers offer little benefit under realistic shift conditions and that recognition‑based defenses can be undermined by attackers who know the model being used.
By Amit Singh Bhatti, Vishal Vaddina
arXiv:2606. 02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat.
By Alexandre Cristov\~ao Maiorano
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness.
arXiv:2608. 04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why.
By Atul Anand, Sourav Chattaraj
arXiv:2609.14976v1 Announce Type: new
Abstract: Long-horizon LLM agents accumulate memory across sessions, creating sparse but high-impact risks: stale facts, conflicting updates, cross-user leakage,...
By Jianhua Jiang, Dongbo Yuan, Weihua Li
SiLR introduces a structure‑preserving admission and process reward mechanism for large language model (LLM) tool agents. Unlike traditional scalar‑score gates that can trap agents in plateau trajectories, SiLR shadow‑executes each proposal and admits it based on a product order over branch‑level violation states, ensuring safe and recoverable actions. Experiments on Gym‑ANM and CityLearn benchmarks show SiLR consistently recovers all multi‑action episodes and outperforms scalar gates, while also providing a robust reward signal for policy learning.
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv:2608. 08029v1 Announce Type: cross Abstract: Khatri et al.
By Alizishaan Khatri, Dun Li Chan
arXiv:2609.13714v1 Announce Type: new
Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...
By Shengwei Zhang, Tao Wu, Fei Qian
Chronicle introduces a method called cut‑point replay to make regression testing of large language model (LLM) agents reproducible. It records an agent’s run at non‑deterministic boundaries as immutable envelopes and then replays selected boundaries while executing the rest live, enabling continuous‑integration tests that detect faulty code changes. Benchmarks show minimal overhead, perfect bit‑stability, and effective detection of unsafe actions in a mutation study.
By Tisha Chawla, Susheem Koul
PentestChain is a ten‑phase automated penetration testing framework that uses a cost‑aware AI cascade, starting with a local 7B‑parameter Ollama model (qwen2.5‑7b) and then free‑tier OpenRouter and Cerebras models, with a rule‑based fallback. It exposes the entire pipeline through a Model Context Protocol (MCP) server that includes eleven tools. The authors evaluate the framework using standard testbeds (AutoPenBench, Cybench subset, PentestGPT 182‑sub‑task benchmark) and report that the local model keeps paid‑API cost at zero while detecting 26 services and enriching 34 CVEs on legacy targets.
By Rushabh Vipulkumar Patel, Dipo Dunsin, Mohammed Almaiah, Mohamed Chahine Ghanem