arXiv AI

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

arXiv:2606. 00152v1 Announce Type: cross Abstract: LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users.

arXiv AI
Sep 17

ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

The paper introduces ASLEval, a framework for measuring privacy exposure displacement in large language model (LLM) agent sessions. It highlights that traditional local proxies—such as inspecting a single action or final response—often miss unauthorized data leaks elsewhere in a multi-step session. ASLEval pre-registers hidden target sets, tracks all declared visible exits, and preserves internal traces for diagnosis, revealing that a single outlet view can overlook nearly 47% of exposure and that internal evidence typically precedes visible leaks. The study underscores the need for benchmarks that define complete visible boundaries, ground claims in pre-specified targets, and report privacy alongside task utility.

By Guosen Wu, Huizhen Huang, Guoxiong Long, Tao Huang, Chen Hou
arXiv AI
Sep 21

CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

CIPL (Channel Inversion for Privacy Leakage) is a channel-aware framework designed to evaluate black-box privacy leakage in large language model agents. It models the leakage process through stages of sensitive source, selection, assembly, execution, observation, and extraction, assessing how selected sensitive units become attacker-recoverable outputs. Experiments across memory, retrieval, and tool-mediated targets, plus a live-agent case study, reveal that recoverability depends on factors beyond storage labels, such as observation surface, prompt alignment, retrieval depth, and provider behavior, and that a semantic audit can uncover disclosures missed by exact matching.

By Tao Huang, Guosen Wu, Guolong Zheng, Jiayang Meng, Chen Hou, Xu Yang, Xuechao Yang, Feng Xia
arXiv Machine Learning
Sep 3

Context Inference Attacks Without Jailbreaks

The paper investigates privacy risks in agentic AI systems that assemble sensitive data into a hidden context before responding. It introduces context‑inference attacks, a security game that evaluates how well attackers can recover this hidden context under varying levels of knowledge and indirect delivery. Experiments show that even with controls such as instructions not to disclose, logit suppression, and context dilution, agents can leak significant contextual information, achieving high success rates across multiple attack settings.

By Prince Jha, Samuele Poppi, Nils Lukas
arXiv Computer Vision
Sep 14

BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines

BodhiPromptShield is a policy‑aware mediation layer for LLM agent pipelines that detects sensitive text spans before they propagate, replacing them with typed placeholders, semantic abstractions, or secure tokens and restoring them only at authorized execution boundaries. In evaluations on AI4Privacy, PrivacyLens, and AgentDojo datasets, the system reduces identifier exposure to 7.4% and 1.8% respectively, and limits exact identifier leakage in final actions to 2.1–3.1%. While mediation preserves factual content according to automated metrics, human annotations show a significant drop in inferability from 100% to 24–53%, indicating the need for human validation of semantic‑leakage measures.

By Bo Ma, Jinsong Wu, Weiqi Yan