The paper introduces the Agent Cross‑Layer Evidence (ACE) corpus, pairing application‑level telemetry with kernel‑level syscall traces to study agent security. It shows that kernel evidence alone is discriminative and that combining it with application‑level data outperforms either layer alone, revealing complementary signals. The study also demonstrates that this cross‑layer approach generalizes to unseen attack families and works across different agent runtimes.
By Spencer King, Zhilu Zhang, Mikhail Kuznetsov, Kay Liu, Baris Coskun, Wei Ding
arXiv:2607. 19742v1 Announce Type: cross Abstract: Cyber Threat Intelligence (CTI) reports richly describe real-world attack processes, but their unstructured narratives cannot be directly used for automated attack-path reasoning.
By Wenbo Hou, Ning Hu, Xueping Wang, Jiahao Gu, Wenjian Luo
arXiv:2608. 15016v1 Announce Type: cross Abstract: Network incident response remains slow and labor-intensive as the defender must infer multi-stage attacks from partial observations and translate recovery decisions into reliable system commands.
By Yiran Gao, Juntao Chen, Tao Li
arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.
By Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
APTInvestBench is a benchmark that evaluates how well autonomous agents can investigate advanced persistent threats (APTs) when faced with different telemetry settings. It contains 370 cases derived from 56 attack reconstructions, totaling 16.4 million log records, and tests agents on seven SOC-inspired telemetry conditions. The benchmark measures evidence acquisition and formal citation support, revealing that while overall coverage drops only slightly when telemetry is limited, a significant portion of actions lose sufficient citation support, highlighting instability in agent performance.
By Yu Wang, Shuhao Li, Tao Yin, Ziyang Li, Xueying Zhao, Peishuai Sun, Jiang Xie
Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response. Y...