arXiv AI By Zixin Rao, Wentian Zhu, Chan Aristella Lu, Zhaorun Chen, Wei Niu, Le Guan, Bo Li, Zhen Xiang

FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and Fusion

Read the original on arXiv AI →

arXiv:2606. 15609v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on long-term memory to support complex task execution, user personalization, and domain adaptation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory

The paper introduces ShadowMem, a defensive framework that protects large language model agents from long-horizon threats by maintaining a dedicated safety-focused memory. Inspired by the shadow stack concept, ShadowMem stores safety-critical context throughout an agent’s execution and uses this shadow memory to evaluate the risk of upcoming actions before they are carried out. Experiments show that ShadowMem outperforms existing defenses in detection accuracy, detects most attacks early, and adds minimal overhead to agent performance.

By Yuhui Wang, Tanqiu Jiang, Jiacheng Liang, Charles Fleming, Ting Wang
arXiv AI
Aug 25

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

InjecMEM introduces a memory injection attack that can steer the responses of large language model agents toward a desired output using only a single interaction, without needing read or edit access to the memory store. The attack leverages the retrieval‑then‑generate workflow of memory systems by crafting a retriever‑agnostic anchor with high‑recall topical cues and an adversarial command optimized through gradient‑based coordinate search. Experiments across various memory systems and backbone models show that InjecMEM reliably induces topic‑conditioned retrieval and targeted generation, remains effective even when memory drifts, and does not affect non‑target queries.

By Hanling Tian, Gengyu Zhang, Zeyang Sha, Jingying Wang, Yuhang Liu, Zhehao Huang, Kun Yang, Xiaolin Huang