APTInvestBench is a benchmark that evaluates how well autonomous agents can investigate advanced persistent threats (APTs) when faced with different telemetry settings. It contains 370 cases derived from 56 attack reconstructions, totaling 16.4 million log records, and tests agents on seven SOC-inspired telemetry conditions. The benchmark measures evidence acquisition and formal citation support, revealing that while overall coverage drops only slightly when telemetry is limited, a significant portion of actions lose sufficient citation support, highlighting instability in agent performance.
By Yu Wang, Shuhao Li, Tao Yin, Ziyang Li, Xueying Zhao, Peishuai Sun, Jiang Xie
arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.
By Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
arXiv:2608. 03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions.
By Xuyang Liu, Yibin Han, Zhenwei Zhang, Kai Chang, Zhiwei Xu, Tian Qiu, Weixian Deng, Jiabao Gao, Xiaolin Peng, Hai Wan, Xibin Zhao
arXiv:2607. 24348v1 Announce Type: cross Abstract: Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature.
By Trung V. Phan, Tri Gia Nguyen, Thomas Bauschert
The paper introduces the Agent Cross‑Layer Evidence (ACE) corpus, pairing application‑level telemetry with kernel‑level syscall traces to study agent security. It shows that kernel evidence alone is discriminative and that combining it with application‑level data outperforms either layer alone, revealing complementary signals. The study also demonstrates that this cross‑layer approach generalizes to unseen attack families and works across different agent runtimes.
By Spencer King, Zhilu Zhang, Mikhail Kuznetsov, Kay Liu, Baris Coskun, Wei Ding
The paper introduces Porting Benchmark, a dataset of 1,234 security patch backporting cases covering cross-version, cross-branch, and cross-repository scenarios, along with a unified evaluation framework. Five tools—spanning program analysis, LLM prompting, and LLM agents—are evaluated under aligned settings, revealing that performance varies significantly across tools and patch complexity, with success rates dropping sharply for structurally complex patches. The study identifies four root-cause categories for failures and demonstrates that reference-based benchmark scores may not fully reflect real-world remediation, as executable validation uncovers additional integration issues.
AgentXploit is a two‑role auditing system that separates repository‑level attack‑path discovery from runtime exploitation for AI agents. The Analyzer Agent traces attacker‑controlled inputs to sensitive operations and records candidate attack paths, while the Exploiter Agent turns these paths into concrete attacks and refines them using runtime feedback. The system is evaluated on AgentXploit‑Bench, a benchmark of 72 reproducible vulnerabilities across 12 open‑source AI‑agent systems, achieving 59.3% end‑to‑end success compared to 38.4% for Codex, and 79.2% attack success on AgentDojo versus 52.7% for AgentVigil.
By Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song
arXiv:2609.15939v1 Announce Type: cross
Abstract: Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can dete...
By Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, Fraser Burch, Takahiro Matsumoto, Jianliang He, Baturay Saglam, Arthur Goldblatt, Zhuoran Yang, Amin Karbasi
arXiv:2606. 02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls.
By Hiskias Dingeto, Will Leeney
arXiv:2606. 18356v1 Announce Type: cross Abstract: Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects.
By Yuchuan Tian, Mengyu Zheng, Haocheng Mei, Ye Yuan, Chao Xu, Xinghao Chen, Hanting Chen, Yu Wang
The paper introduces Porting Benchmark, a curated dataset of 1,234 security patch backporting cases that span cross-version, cross-branch, and cross-repository scenarios, along with a common evaluation framework. Five tools—spanning program analysis, LLM prompting, and LLM agents—are evaluated under aligned settings, revealing that performance varies significantly across tools and that complex patches (Type-IV) see a sharp drop in success rate. The study identifies four root-cause categories for failures and demonstrates that reference-based benchmark scores may not fully capture real-world remediation, as executable validation uncovers additional integration issues.
By Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang, Fangyuan Zhang, Bingyang Ren, Yang Liu, Hui Li
Large Language Models (LLMs) have shown promise for automated penetration testing, yet existing end-to-end black-box evaluations are highly susceptible to error cascading: failures in early reconnaissance can mask an agent's actual ability to exploit vulnerabilities. To more accurately characterize these capabilities, we propose a two-stage decoupled evaluation framework that separates exploit execution from reconnaissance.