arXiv AI

AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems

arXiv:2602. 11510v3 Announce Type: replace Abstract: Multi-agent Large Language Model (LLM) systems create privacy risks that current output-only benchmarks cannot measure.

arXiv AI
Sep 21

CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

CIPL (Channel Inversion for Privacy Leakage) is a channel-aware framework designed to evaluate black-box privacy leakage in large language model agents. It models the leakage process through stages of sensitive source, selection, assembly, execution, observation, and extraction, assessing how selected sensitive units become attacker-recoverable outputs. Experiments across memory, retrieval, and tool-mediated targets, plus a live-agent case study, reveal that recoverability depends on factors beyond storage labels, such as observation surface, prompt alignment, retrieval depth, and provider behavior, and that a semantic audit can uncover disclosures missed by exact matching.

By Tao Huang, Guosen Wu, Guolong Zheng, Jiayang Meng, Chen Hou, Xu Yang, Xuechao Yang, Feng Xia
arXiv Machine Learning
Sep 7

Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys

The paper reports a privacy breach in a two-node split‑LLM training system where the returned gradient reveals which data rows were real, despite the system passing standard privacy checks. By exploiting the fact that decoy rows produce zero gradients, an attacker can identify real rows with 100% accuracy across multiple runs. The authors demonstrate that adding gradient clipping and noise can mitigate the leak, but the system remains vulnerable to several untested attack vectors.

By Georgios Politis, Evangelos Pappas