PentestChain is a ten‑phase automated penetration testing framework that uses a cost‑aware AI cascade, starting with a local 7B‑parameter Ollama model (qwen2.5‑7b) and then free‑tier OpenRouter and Cerebras models, with a rule‑based fallback. It exposes the entire pipeline through a Model Context Protocol (MCP) server that includes eleven tools. The authors evaluate the framework using standard testbeds (AutoPenBench, Cybench subset, PentestGPT 182‑sub‑task benchmark) and report that the local model keeps paid‑API cost at zero while detecting 26 services and enriching 34 CVEs on legacy targets.
By Rushabh Vipulkumar Patel, Dipo Dunsin, Mohammed Almaiah, Mohamed Chahine Ghanem
arXiv:2607. 13085v1 Announce Type: cross Abstract: Recent autonomous penetration testing papers report high benchmark scores while adding multi-component security harnesses around frontier LLMs.
By Ananda Dhakal, Krish Neupane, Aarjan Chaudhary
arXiv:2605. 23243v2 Announce Type: replace-cross Abstract: We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/Java/Python) and black-box web application security testing (five production-style applications with 118 ground-truth vulnerabilities across 20+ CWE families, which we will open-source).
By Vivek Dahiya, Sunny Nehra, Vipul Dholariya, Bhavik Shangari, Chandra Khatri
arXiv:2609.22664v1 Announce Type: cross
Abstract: Research on large language model agents for penetration testing is evaluated almost entirely by capability: whether the agent captures a flag or repr...
By Joas Antonio dos Santos Barbosa
arXiv:2608.21423v1 Announce Type: cross
Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to...
By Israt Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Jakir Hossain, M. F. Mridha
arXiv:2605. 10834v2 Announce Type: replace Abstract: AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets.
By Pedro Conde, Henrique Branquinho, Valerio Mazzone, Bruno Mendes, Andr\'e Baptista, Nuno Moniz
arXiv:2603. 19864v2 Announce Type: replace Abstract: Penetration testing, the practice of simulating cyberattacks to identify vulnerabilities, is a complex sequential decision-making task that is inherently partially observable and features large action spaces.
By Raphael Simon, Jos\'e Carrasquel, Wim Mees, Pieter Libin
arXiv:2605. 09504v2 Announce Type: replace-cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, parallel exploration, and evolutionary optimization.
By Michael A. Riegler, Inga Str\"umke
arXiv:2605. 26548v2 Announce Type: replace-cross Abstract: Finding a real vulnerability in complicated systems is a challenging, long-horizon task that demands reasoning across an entire codebase to produce a working proof-of-concept (PoC).
By Hwiwon Lee, Jiawei Liu, Dongjun Kim, Wubing Xia, Ziqi Zhang, Chunqiu Steven Xia, Lingming Zhang
End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where. We present layer-isolated evaluation: a deployed ordering agent is decomposed into a fixed taxonomy of layers (ontology, intent, routing, decomposition, escalation, safety, memory, and cross-cutting envelope/defense), each exercised by its own assertion slice in a deterministic, no-LLM "pure" mode.
arXiv:2605.02346v2 Announce Type: replace-cross
Abstract: Operational technology (OT) devices run safety-critical physical processes, yet their security testing remains manual and expert-intensive. A...
By Adel ElZemity, Budi Arief, Shujun Li, George Oikonomou, James Pope
The paper proposes using lightweight, calibrated System One decision models—specifically JEV and Laya—to improve autonomous penetration-testing harnesses that rely on large language models (LLMs). It defines four key decision points (finding adjudication, severity recalibration, agent pruning, and confirmation loops) and presents a NeuroSploit case study showing differences in severity distribution, runtime, and grading when using TypeSafe System One. The authors review existing System One specifications, discuss various RL-based training approaches, and introduce Rave, a domain‑adapted model with a proposed training and evaluation framework.
By Joas Antonio dos Santos Barbosa