APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
APTInvestBench is a benchmark that evaluates how well autonomous agents can investigate advanced persistent threats (APTs) when faced with different telemetry settings. It contains 370 cases derived from 56 attack reconstructions, totaling 16.4 million log records, and tests agents on seven SOC-inspired telemetry conditions. The benchmark measures evidence acquisition and formal citation support, revealing that while overall coverage drops only slightly when telemetry is limited, a significant portion of actions lose sufficient citation support, highlighting instability in agent performance.
arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.
arXiv:2608. 03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions.
arXiv:2607. 24348v1 Announce Type: cross Abstract: Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature.
The paper introduces the Agent Cross‑Layer Evidence (ACE) corpus, pairing application‑level telemetry with kernel‑level syscall traces to study agent security. It shows that kernel evidence alone is discriminative and that combining it with application‑level data outperforms either layer alone, revealing complementary signals. The study also demonstrates that this cross‑layer approach generalizes to unseen attack families and works across different agent runtimes.
The paper introduces Porting Benchmark, a dataset of 1,234 security patch backporting cases covering cross-version, cross-branch, and cross-repository scenarios, along with a unified evaluation framework. Five tools—spanning program analysis, LLM prompting, and LLM agents—are evaluated under aligned settings, revealing that performance varies significantly across tools and patch complexity, with success rates dropping sharply for structurally complex patches. The study identifies four root-cause categories for failures and demonstrates that reference-based benchmark scores may not fully reflect real-world remediation, as executable validation uncovers additional integration issues.