arXiv AI

The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents

arXiv AI
Jul 29

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

arXiv:2607. 24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run.

By Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu
arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
arXiv AI
Sep 10

Structurally Close, Temporally Distant: Measuring Security Exposure in Long-Horizon LLM Agents

The paper introduces a provenance‑aware execution graph for long‑horizon LLM agents, defining influence distance (DI) as the shortest structural path from an untrusted source to a sensitive action. Compared to the traditional sequence distance (DT), DI is always less than or equal to DT, revealing a median gap of nine hops in 454 injection–sink pairs across multiple models and datasets. The study shows that most pairs exhibit a non‑zero gap, and a deterministic DI‑based gate can block attacks missed by a sequence‑only gate without extra benign blocking.

By Md Jafrin Hossain, Nur Al Hasan Haldar