arXiv AI By Tao Huang, Guosen Wu, Guolong Zheng, Jiayang Meng, Chen Hou, Xu Yang, Xuechao Yang, Feng Xia

CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

Read the original on arXiv AI →

CIPL (Channel Inversion for Privacy Leakage) is a channel-aware framework designed to evaluate black-box privacy leakage in large language model agents. It models the leakage process through stages of sensitive source, selection, assembly, execution, observation, and extraction, assessing how selected sensitive units become attacker-recoverable outputs. Experiments across memory, retrieval, and tool-mediated targets, plus a live-agent case study, reveal that recoverability depends on factors beyond storage labels, such as observation surface, prompt alignment, retrieval depth, and provider behavior, and that a semantic audit can uncover disclosures missed by exact matching.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 3

MNC: Scope-Bound Semantic Declassification for Private LLM-Agent Communication

Multi-agent large language model (LLM) systems can expose protected state through internal messages, tool arguments, logs, and persistent memory even when their public outputs appear innocuous. Existing privacy prompts, redaction methods, and source-level access controls restrict surface content or data access, but do not specify what a legitimately informed agent should disclose or how that disclosure may be reused downstream.

arXiv AI
Sep 17

ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

The paper introduces ASLEval, a framework for measuring privacy exposure displacement in large language model (LLM) agent sessions. It highlights that traditional local proxies—such as inspecting a single action or final response—often miss unauthorized data leaks elsewhere in a multi-step session. ASLEval pre-registers hidden target sets, tracks all declared visible exits, and preserves internal traces for diagnosis, revealing that a single outlet view can overlook nearly 47% of exposure and that internal evidence typically precedes visible leaks. The study underscores the need for benchmarks that define complete visible boundaries, ground claims in pre-specified targets, and report privacy alongside task utility.

By Guosen Wu, Huizhen Huang, Guoxiong Long, Tao Huang, Chen Hou