arXiv:2606. 11686v1 Announce Type: cross Abstract: End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where.
By Sawyer Zhang, Alexander Wang, Sophie Lei
FaultLens is a method for creating compact behavioral test suites for generated operational programs, balancing thoroughness with cost. It executes a rich probe domain once, stores fault‑probe kill relations, and learns probe orderings from earlier program generations using a fault‑driven greedy component and a mutation‑independent diversity component. In evaluations across multiple environments and program generations, a 32‑probe hybrid suite achieved 99.0% coverage of dynamically killable faults while using only 1.2‑2.0% of the exhaustive domain, and improved macro coverage when a fault family was withheld from training.
FaultLens is a technique for generating compact behavioral test suites for programs produced by automated generators. It learns probe orderings from earlier program generations, combining a fault‑driven greedy component with a mutation‑independent diversity component to cover a wide range of probe families, cases, templates, and temporal bins. In experiments on twenty generated operational policies across four environments, a 32‑probe hybrid suite learned from early generations covered 99.0% of dynamically killable faults in later generations while using only 1.2–2.0% of the exhaustive test domain.
By Zeming Liu, Hang Lyu, Jingtao Zhang
PentestChain is a ten‑phase automated penetration testing framework that uses a cost‑aware AI cascade, starting with a local 7B‑parameter Ollama model (qwen2.5‑7b) and then free‑tier OpenRouter and Cerebras models, with a rule‑based fallback. It exposes the entire pipeline through a Model Context Protocol (MCP) server that includes eleven tools. The authors evaluate the framework using standard testbeds (AutoPenBench, Cybench subset, PentestGPT 182‑sub‑task benchmark) and report that the local model keeps paid‑API cost at zero while detecting 26 services and enriching 34 CVEs on legacy targets.
By Rushabh Vipulkumar Patel, Dipo Dunsin, Mohammed Almaiah, Mohamed Chahine Ghanem
arXiv:2609.14976v1 Announce Type: new
Abstract: Long-horizon LLM agents accumulate memory across sessions, creating sparse but high-impact risks: stale facts, conflicting updates, cross-user leakage,...
By Jianhua Jiang, Dongbo Yuan, Weihua Li
arXiv:2609.35889v1 Announce Type: cross
Abstract: Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing...
By Xiaoyu Xu, Zi Liang, Minxin Du, Qipeng Xie, Qingqing Ye, Yuyuan Li, Haibo Hu
Chronicle introduces a method called cut‑point replay to make regression testing of large language model (LLM) agents reproducible. It records an agent’s run at non‑deterministic boundaries as immutable envelopes and then replays selected boundaries while executing the rest live, enabling continuous‑integration tests that detect faulty code changes. Benchmarks show minimal overhead, perfect bit‑stability, and effective detection of unsafe actions in a mutation study.
By Tisha Chawla, Susheem Koul
arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.
By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness.
arXiv:2609.13714v1 Announce Type: new
Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...
By Shengwei Zhang, Tao Wu, Fei Qian
AI-driven penetration testing has been demonstrated with premium frontier models such as GPT-4, but the per-engagement token cost makes continuous, automated testing unaffordable for the smaller organ...
arXiv:2607. 16345v1 Announce Type: cross Abstract: Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task.
By Tejas Singh Anand, Yuet Ying Christina Wang, Wanting Jiang, Steve Masson, Tian Zheng, Bingjie Zhou