Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response. Y...
arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.
By Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
arXiv:2608. 03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions.
By Xuyang Liu, Yibin Han, Zhenwei Zhang, Kai Chang, Zhiwei Xu, Tian Qiu, Weixian Deng, Jiabao Gao, Xiaolin Peng, Hai Wan, Xibin Zhao
The paper introduces the Agent Cross‑Layer Evidence (ACE) corpus, pairing application‑level telemetry with kernel‑level syscall traces to study agent security. It shows that kernel evidence alone is discriminative and that combining it with application‑level data outperforms either layer alone, revealing complementary signals. The study also demonstrates that this cross‑layer approach generalizes to unseen attack families and works across different agent runtimes.
By Spencer King, Zhilu Zhang, Mikhail Kuznetsov, Kay Liu, Baris Coskun, Wei Ding
AgentXploit is a two‑role auditing system that separates repository‑level attack‑path discovery from runtime exploitation for AI agents. The Analyzer Agent traces attacker‑controlled inputs to sensitive operations and records candidate attack paths, while the Exploiter Agent turns these paths into concrete attacks and refines them using runtime feedback. The system is evaluated on AgentXploit‑Bench, a benchmark of 72 reproducible vulnerabilities across 12 open‑source AI‑agent systems, achieving 59.3% end‑to‑end success compared to 38.4% for Codex, and 79.2% attack success on AgentDojo versus 52.7% for AgentVigil.
By Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song
The paper introduces Porting Benchmark, a curated dataset of 1,234 security patch backporting cases that span cross-version, cross-branch, and cross-repository scenarios, along with a common evaluation framework. Five tools—spanning program analysis, LLM prompting, and LLM agents—are evaluated under aligned settings, revealing that performance varies significantly across tools and that complex patches (Type-IV) see a sharp drop in success rate. The study identifies four root-cause categories for failures and demonstrates that reference-based benchmark scores may not fully capture real-world remediation, as executable validation uncovers additional integration issues.
By Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang, Fangyuan Zhang, Bingyang Ren, Yang Liu, Hui Li
The paper introduces Porting Benchmark, a dataset of 1,234 security patch backporting cases covering cross-version, cross-branch, and cross-repository scenarios, along with a unified evaluation framework. Five tools—spanning program analysis, LLM prompting, and LLM agents—are evaluated under aligned settings, revealing that performance varies significantly across tools and patch complexity, with success rates dropping sharply for structurally complex patches. The study identifies four root-cause categories for failures and demonstrates that reference-based benchmark scores may not fully reflect real-world remediation, as executable validation uncovers additional integration issues.
arXiv:2609.15939v1 Announce Type: cross
Abstract: Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can dete...
By Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, Fraser Burch, Takahiro Matsumoto, Jianliang He, Baturay Saglam, Arthur Goldblatt, Zhuoran Yang, Amin Karbasi
arXiv:2607. 24348v1 Announce Type: cross Abstract: Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature.
By Trung V. Phan, Tri Gia Nguyen, Thomas Bauschert
arXiv:2606. 02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls.
By Hiskias Dingeto, Will Leeney
arXiv:2607. 14570v1 Announce Type: new Abstract: AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an agent completes its assigned task while covertly weakening safeguards through actions such as broadening permissions, degrading logging, or introducing persistence mechanisms.
By Preeti Ravindra, Rahul Tiwari, Vincent Wolowski
arXiv:2609.14593v1 Announce Type: cross
Abstract: Living-Off-the-Land (LOTL) is the dominant evasion technique of Advanced Persistent Threat (APT) actors, exploiting legitimate Windows utilities to c...
By Ahad Bin Islam Shoeb, Kamrul Hasan, Jamal Uddin Tanvin, Liang Hong, Imtiaz Ahmed, Md Arif Billah, Al Amin