arXiv:2606. 17464v1 Announce Type: new Abstract: Membership inference attacks (MIAs) are a canonical way to assess a machine learning model's privacy properties.
By Jeffrey G. Wang, Jason Wang, Marvin Li, Seth Neel
The paper examines the reliability of membership inference attack (MIA) vulnerability evaluation. It identifies two weaknesses: finite‑sample bias from sampling shadow datasets from a fixed superset, and miscalibration when aggregating true positive rates across individuals at very low false positive rates. The authors propose simple fixes that avoid extra computational cost and suggest further improvements with additional computation.
By Joonas J\"alk\"o, Gauri Pradhan, Ossi R\"ais\"a, Antti Honkela
arXiv:2606. 14210v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in privacy-sensitive domains, where users must balance the risk of data exposure through external APIs against the high computational cost of local deployment.
By Zixuan Gu, Xiaojun Ye, Yang Liu
arXiv:2606. 17110v1 Announce Type: cross Abstract: Large Language Models are increasingly trained on proprietary or sensitive data, from private healthcare and financial records to user conversations containing secrets.
By Md Abdullah Al Mamun, Ngoc Phu Doan, Pedram Zaree, Ihsen Alouani, Nael Abu-Ghazaleh
UniGuardian is a training‑free detector for large language models that jointly identifies prompt injection, backdoor, and adversarial attacks—collectively called Prompt Trigger Attacks (PTA). It measures how structured prompt perturbations shift the model’s output distribution and uses a single‑forward strategy to detect attacks while generating text in a shared batched forward pass. Experiments show that UniGuardian accurately and efficiently identifies trigger‑activated prompts in LLMs.
By Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao
arXiv:2606. 31991v1 Announce Type: cross Abstract: The tendency of large generative models to memorize training data makes sample verification critical for privacy auditing and copyright enforcement.
By Wojciech {\L}apacz, Stanis{\l}aw Pawlak