arXiv:2605. 08442v5 Announce Type: replace-cross Abstract: We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not necessarily block execution, and vice versa.
By Jun Wen Leong
The paper introduces a provenance‑aware execution graph for long‑horizon LLM agents, defining influence distance (DI) as the shortest structural path from an untrusted source to a sensitive action. Compared to the traditional sequence distance (DT), DI is always less than or equal to DT, revealing a median gap of nine hops in 454 injection–sink pairs across multiple models and datasets. The study shows that most pairs exhibit a non‑zero gap, and a deterministic DI‑based gate can block attacks missed by a sequence‑only gate without extra benign blocking.
By Md Jafrin Hossain, Nur Al Hasan Haldar
arXiv:2606. 25115v1 Announce Type: new Abstract: On-device language-model agents improve by accumulating experience in retrieved memory rather than by updating weights.
By Beining Wu, Zihao Ding, Jun Huang, Yanxiao Zhao
arXiv:2605. 08442v3 Announce Type: replace-cross Abstract: Persistent memory attacks against LLM agents achieve high attack success rates against open-source models.
By Jun Wen Leong
arXiv:2606. 30602v1 Announce Type: cross Abstract: Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows.
By Kunyang Li, Kyle Domico, Jonathan Gregory, Patrick McDaniel
arXiv:2602. 11510v3 Announce Type: replace Abstract: Multi-agent Large Language Model (LLM) systems create privacy risks that current output-only benchmarks cannot measure.
By Faouzi El Yagoubi, Godwin Badu-Marfo, Ranwa Al Mallah
arXiv:2606. 15903v1 Announce Type: cross Abstract: Where an LLM sits in an agent memory pipeline -- between the recall plane that retrieves stored facts (extensively benchmarked) and the control plane that mutates them via supersede, release, purge (largely untested) -- shapes which forgetting failure modes the system recovers.
By Dongxu Yang
CIPL (Channel Inversion for Privacy Leakage) is a channel-aware framework designed to evaluate black-box privacy leakage in large language model agents. It models the leakage process through stages of sensitive source, selection, assembly, execution, observation, and extraction, assessing how selected sensitive units become attacker-recoverable outputs. Experiments across memory, retrieval, and tool-mediated targets, plus a live-agent case study, reveal that recoverability depends on factors beyond storage labels, such as observation surface, prompt alignment, retrieval depth, and provider behavior, and that a semantic audit can uncover disclosures missed by exact matching.
By Tao Huang, Guosen Wu, Guolong Zheng, Jiayang Meng, Chen Hou, Xu Yang, Xuechao Yang, Feng Xia
arXiv:2609.37367v1 Announce Type: cross
Abstract: Decentralized large language model (LLM) fine-tuning lets organizations collaboratively train a shared LLM on data they cannot pool, without a centra...
By Sayan Biswas, Jade Garcia Bourr\'ee, Rachid Guerraoui, Maxime Jacovella, Anne-Marie Kermarrec, Sathwika Peechara, Martijn de Vos, Milos Vujasinovic
arXiv:2608. 08131v1 Announce Type: cross Abstract: In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activates the concealed condition, and protective authority turns against the system.
By Satoshi Matsuoka
arXiv:2609.37567v1 Announce Type: cross
Abstract: Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collabor...
By Longzhu He, Zelang Wen, Xinfeng Li, Sen Su, XiaoFeng Wang
arXiv:2606. 30566v1 Announce Type: cross Abstract: We discover a behavioral invariant in LLM agents under persistent memory poisoning: in architectures where routing information is retrieved through observable memory-tool invocations, successful attacks require calling memory_recall_fact before email_send_email, a transition that non-exfiltrating sessions rarely exhibit.
By Jun Wen Leong