arXiv AI

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans

arXiv:2606. 09844v1 Announce Type: cross Abstract: Large Language Models (LLMs) alter their privacy behavior based on the perceived identity of their interlocutor.

arXiv AI
Aug 11

Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents

arXiv:2510. 04465v3 Announce Type: replace-cross Abstract: LLM agents require personal information for personalization in order to effectively act on users' behalf, but this raises privacy concerns that can discourage data sharing, limiting both the autonomy levels at which agents can operate and the effectiveness of personalization.

By Zhiping Zhang, Yi Evie Zhang, Freda Shi, Tianshi Li
arXiv AI
Sep 25

PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations

PrivDrift is a benchmark that tests whether user‑disclosed secrets can still be recovered by large language models after the conversation shifts to unrelated topics. It includes 1,000 controlled multi‑turn dialogues with seeded secrets, topic‑drift turns, and standardized extraction probes. Experiments on three LLMs with extended context windows show that dialogue‑level leakage remains substantial—between 38.7% and 54.6%—and is influenced by model, secret type, and persuasion intensity, while additional topic drift does not reliably reduce leakage.

By Luciano Maldonado
arXiv Computer Vision
Sep 14

BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines

BodhiPromptShield is a policy‑aware mediation layer for LLM agent pipelines that detects sensitive text spans before they propagate, replacing them with typed placeholders, semantic abstractions, or secure tokens and restoring them only at authorized execution boundaries. In evaluations on AI4Privacy, PrivacyLens, and AgentDojo datasets, the system reduces identifier exposure to 7.4% and 1.8% respectively, and limits exact identifier leakage in final actions to 2.1–3.1%. While mediation preserves factual content according to automated metrics, human annotations show a significant drop in inferability from 100% to 24–53%, indicating the need for human validation of semantic‑leakage measures.

By Bo Ma, Jinsong Wu, Weiqi Yan
arXiv AI
Aug 24

Personalized Privacy Control in LLMs via Attention Head Intervention

The paper introduces the concept of personalized privacy for large language models, allowing user‑specific disclosure preferences to guide information sharing. It presents P3Bench, a benchmark that extends contextual privacy policies with personalized rules, and shows that existing prompt‑based methods often ignore these policies. To improve compliance, the authors propose “Repair”, an inference‑time attention head intervention that aligns model responses with user‑specific privacy rules.

By Junseok Kim, Nakyeong Yang, Kyomin Jung
Hugging Face Trending Papers
Aug 20

Inadvertent Context Leakage in Language Models

For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction.