arXiv AI

Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents

arXiv:2608. 05199v1 Announce Type: cross Abstract: Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique.

arXiv AI
Sep 10

Structurally Close, Temporally Distant: Measuring Security Exposure in Long-Horizon LLM Agents

The paper introduces a provenance‑aware execution graph for long‑horizon LLM agents, defining influence distance (DI) as the shortest structural path from an untrusted source to a sensitive action. Compared to the traditional sequence distance (DT), DI is always less than or equal to DT, revealing a median gap of nine hops in 454 injection–sink pairs across multiple models and datasets. The study shows that most pairs exhibit a non‑zero gap, and a deterministic DI‑based gate can block attacks missed by a sequence‑only gate without extra benign blocking.

By Md Jafrin Hossain, Nur Al Hasan Haldar
arXiv Machine Learning
Sep 11

DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents

DriftNet is a dual‑head trajectory Transformer designed to detect and localize prompt injection attacks in large language model agents. It processes logged tool‑call trajectories, classifying each as compromised or not while labeling every step as benign, injection point, hijacked, or failed injection. On the AgentDrift benchmark, DriftNet achieves high accuracy, with an F1 score of 0.983, 98.7% exact injection‑point recovery, and low false‑alarm rates.

By Asif Pinjari, Mithun Paul Saint-Germain
arXiv Machine Learning
Sep 7

Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection

The paper investigates how to properly validate candidate models before promoting them to replace incumbent classifiers in adaptive network intrusion detection systems. It demonstrates that promotion decisions can be biased by how challengers are constructed and the amount of evidence they receive, and that using self‑contained challenger pipelines and sufficient candidate evidence reduces apparent promotion harm. The study also shows that policy rankings shift with candidate comparability and that no single update policy dominates across benchmarks.

By Roberto Fern\'andez-Barrios, Iker Pastor-L\'opez, Amaia Pikatza-Huerga, Pablo Garc\'ia Bringas
arXiv AI
Jun 4

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

arXiv:2602. 06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks.

By Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee, Matthew Kowal, Nayeema Nonta, Samuel Simko, Stephen Casper, Zhijing Jin, Kellin Pelrine, Sirisha Rambhatla