arXiv Machine Learning

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

arXiv:2608. 03130v1 Announce Type: cross Abstract: Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attributes even when they are never stated explicitly.

arXiv AI
2d ago

Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs

Tokenized Key-Gated Adapter Routing (Locket) is a framework that embeds fine‑grained, policy‑driven access control into large language models by training lightweight LoRA adapters for different privacy policies. A gating module associates a learned keyed entry token with a specific adapter, allowing authorized tokens to unlock private knowledge while invalid or missing tokens trigger privacy‑preserving adapters that redact or sanitize sensitive content. Experiments on datasets such as Enron, ECHR, and Yelp with models like Qwen3, Llama‑3.2, and Gemma‑2‑2B show that Locket maintains perplexity comparable to fine‑tuning when the correct token is provided, and significantly reduces PII leakage when the token is absent or invalid, without sacrificing utility.

By Mohamed Shaaban, Mohamed Elmahallawy
Hugging Face Trending Papers
Aug 3

MNC: Scope-Bound Semantic Declassification for Private LLM-Agent Communication

Multi-agent large language model (LLM) systems can expose protected state through internal messages, tool arguments, logs, and persistent memory even when their public outputs appear innocuous. Existing privacy prompts, redaction methods, and source-level access controls restrict surface content or data access, but do not specify what a legitimately informed agent should disclose or how that disclosure may be reused downstream.

arXiv AI
Sep 21

CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

CIPL (Channel Inversion for Privacy Leakage) is a channel-aware framework designed to evaluate black-box privacy leakage in large language model agents. It models the leakage process through stages of sensitive source, selection, assembly, execution, observation, and extraction, assessing how selected sensitive units become attacker-recoverable outputs. Experiments across memory, retrieval, and tool-mediated targets, plus a live-agent case study, reveal that recoverability depends on factors beyond storage labels, such as observation surface, prompt alignment, retrieval depth, and provider behavior, and that a semantic audit can uncover disclosures missed by exact matching.

By Tao Huang, Guosen Wu, Guolong Zheng, Jiayang Meng, Chen Hou, Xu Yang, Xuechao Yang, Feng Xia
arXiv AI
4d ago

Audience-Bound Persistent Memory: Authorization Across the Memory Lifecycle

The paper introduces Audience‑Bound Persistent Memory, a system that tracks the audience of each memory item and enforces authorization throughout the memory lifecycle. Each item carries the audience present at recording, and derived items are partitioned or suppressed based on the intersection of source audiences, expanding only through explicit grants. The authors implement the approach in two reference architectures—a flat store and a relationship graph—and evaluate it on 10,000 multi‑party histories, showing that no forbidden items entered any context while unscoped retrieval exposed forbidden items in 82% of cases, and that entitled recall matched policy‑equivalent baselines and outperformed unscoped retrieval by 0.30 Recall@5.

By Sibo Liu
arXiv Machine Learning
Sep 3

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

The paper introduces CAPTURE, a system designed to help personalized language agents distinguish genuine preference changes from temporary context shifts or malicious memory poisoning. CAPTURE employs a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing to resolve ambiguity. Experiments on 480 episodes from 96 users show CAPTURE outperforms baseline methods, limiting poisoning success while accepting most real preference updates.

By S M Asif Hossain, Ruksat Khan Shayoni, Md Kishor Morol