arXiv Computation and Language

LLM Anonymization Against Agentic Re-Identification

The paper introduces AURA, an LLM-powered mask‑reconstruct framework for anonymizing text while preserving utility. It decouples privacy localization from utility‑preserving reconstruction and uses adversarial checks to select candidates. Experiments on real‑user interview transcripts show that AURA achieves the lowest agentic re‑identification rates among non‑DP methods and retains more contextual utility than prior LLM anonymizers.

arXiv AI
Sep 12

Demystifying the Privacy-Utility Trade-off in LLM Interactions

The paper investigates how privacy-preserving sanitization of user context in large language model (LLM) interactions affects downstream performance. It identifies three mechanisms—Context‑Dependent Utility, Strategic Adaptation, and Combinatorial Interplay—that explain when and how to sanitize data. Based on these insights, the authors propose an intent‑driven local protection framework using a lightweight model (Veilmind‑4B) to dynamically extract, sanitize, and restore context, achieving lower privacy leakage while maintaining higher utility than existing baselines.

By Zhenhua Liu, Zhanxu Xie, Junjie Yu, Tong Zhu, Lijun Li, Wenliang Chen
arXiv Machine Learning
Aug 5

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

arXiv:2608. 03130v1 Announce Type: cross Abstract: Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attributes even when they are never stated explicitly.

By Jong Wook Kim, Byoungjae Min, Kennedy Edemacu, Yoonhyuk Choi, Sae-Hong Cho, Beakcheol Jang
arXiv Machine Learning
Sep 3

Context Inference Attacks Without Jailbreaks

The paper investigates privacy risks in agentic AI systems that assemble sensitive data into a hidden context before responding. It introduces context‑inference attacks, a security game that evaluates how well attackers can recover this hidden context under varying levels of knowledge and indirect delivery. Experiments show that even with controls such as instructions not to disclose, logit suppression, and context dilution, agents can leak significant contextual information, achieving high success rates across multiple attack settings.

By Prince Jha, Samuele Poppi, Nils Lukas
arXiv Machine Learning
Oct 2

SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs

The paper introduces SparLeak, a side‑channel attack that exploits a new GPU micro‑architectural leakage called Sparsity‑Induced Memory Access (SIMA) caused by sparse attention in large language models. By capturing SIMA traces during LLM inference, SparLeak can infer user‑query attributes from prefill‑phase traces and reconstruct autoregressive responses from decoding‑phase traces. Experiments on three LLM architectures, three sparse attention mechanisms, and three privacy‑sensitive datasets show high success rates (90.9% for attribute inference and 87.3% for response reconstruction), underscoring the need to address SIMA leakage in sparse‑attention deployments.

By Fahao Chen, Linkang Du, Jinhao Zhou, Peng Li, Zhou Su
arXiv Computation and Language
Sep 11

On the Impact of Anonymization on the Performance of Large Language Models

The paper systematically studies how anonymizing input data affects large language models (LLMs). Five prominent LLMs were evaluated on eleven benchmarks, comparing performance on original versus pseudonymized inputs. Results show that anonymization generally degrades performance, with larger drops for more capable models and task-dependent effects; reversible anonymization preserves entity uniqueness better than irreversible redaction, and prompting about anonymization offers no benefit.

By Tobias Deu{\ss}er, Max Hahnb\"uck, Lorenz Sparrenberg, Tobias Uelwer, Christian Bauckhage, Rafet Sifa
arXiv Computer Vision
Sep 14

BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines

BodhiPromptShield is a policy‑aware mediation layer for LLM agent pipelines that detects sensitive text spans before they propagate, replacing them with typed placeholders, semantic abstractions, or secure tokens and restoring them only at authorized execution boundaries. In evaluations on AI4Privacy, PrivacyLens, and AgentDojo datasets, the system reduces identifier exposure to 7.4% and 1.8% respectively, and limits exact identifier leakage in final actions to 2.1–3.1%. While mediation preserves factual content according to automated metrics, human annotations show a significant drop in inferability from 100% to 24–53%, indicating the need for human validation of semantic‑leakage measures.

By Bo Ma, Jinsong Wu, Weiqi Yan