From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2607. 27080v1 Announce Type: cross Abstract: Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist.
The paper introduces CAPTURE, a system designed to help personalized language agents distinguish genuine preference changes from temporary context shifts or malicious memory poisoning. CAPTURE employs a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing to resolve ambiguity. Experiments on 480 episodes from 96 users show CAPTURE outperforms baseline methods, limiting poisoning success while accepting most real preference updates.
arXiv:2605.01970v4 Announce Type: replace-cross Abstract: Memory systems enable otherwise stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. Th...
Memory is becoming a default subsystem in deployed LLM agents to provide persistent personalization and continuity. This naturally prompts a question: will memory system introduce new vulnerabilities...
arXiv:2606. 06054v1 Announce Type: new Abstract: Personal AI agents increasingly rely on long-term memory to provide persistent personalization across sessions.
arXiv:2605. 08442v3 Announce Type: replace-cross Abstract: Persistent memory attacks against LLM agents achieve high attack success rates against open-source models.