arXiv Machine Learning

Automated Membership Inference Attacks (AutoMIA): Discovering MIA Signal Computations using LLM Agents

The paper introduces AutoMIA, a framework that uses large language model agents to automatically design and implement new membership inference attack (MIA) signal computations. By systematically exploring a wide range of attack strategies, AutoMIA discovers novel MIAs tailored to specific target models and datasets, achieving up to a 0.18 absolute improvement in AUC over existing methods. This demonstrates that LLM agents can serve as an effective and scalable approach for creating state‑of‑the‑art MIAs.

arXiv Machine Learning
1d ago

On Reliability of Membership Inference Vulnerability Evaluation

The paper examines the reliability of membership inference attack (MIA) vulnerability evaluation. It identifies two weaknesses: finite‑sample bias from sampling shadow datasets from a fixed superset, and miscalibration when aggregating true positive rates across individuals at very low false positive rates. The authors propose simple fixes that avoid extra computational cost and suggest further improvements with additional computation.

By Joonas J\"alk\"o, Gauri Pradhan, Ossi R\"ais\"a, Antti Honkela
arXiv AI
2d ago

UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

UniGuardian is a training‑free detector for large language models that jointly identifies prompt injection, backdoor, and adversarial attacks—collectively called Prompt Trigger Attacks (PTA). It measures how structured prompt perturbations shift the model’s output distribution and uses a single‑forward strategy to detect attacks while generating text in a shared batched forward pass. Experiments show that UniGuardian accurately and efficiently identifies trigger‑activated prompts in LLMs.

By Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao
arXiv Machine Learning
Sep 3

Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models

The paper introduces a black-box membership inference attack framework tailored for fine-tuned text-to-speech models, addressing challenges in query generation and representation engineering. It evaluates five query types, finding recitation queries most effective, and uses multi-level speech embeddings with temporal alignment for fine-grained comparison. Experiments on CosyVoice2, F5-TTS, and XTTS-v2 trained on VCTK and British Dialect datasets show high privacy leakage, with speaker-level AUC above 0.80 and record-level AUC between 0.80 and 0.90.

By Kunlin Cai, Kaiyuan Zhang, Zihang Xiang, Jinghuai Zhang, Abeer Alwan, Fnu Suya, Yuan Tian