arXiv Machine Learning By Toan Tran, Olivera Kotevska, Li Xiong

Automated Membership Inference Attacks (AutoMIA): Discovering MIA Signal Computations using LLM Agents

Read the original on arXiv Machine Learning →

The paper introduces AutoMIA, a framework that uses large language model agents to automatically design and implement new membership inference attack (MIA) signal computations. By systematically exploring a wide range of attack strategies, AutoMIA discovers novel MIAs tailored to specific target models and datasets, achieving up to a 0.18 absolute improvement in AUC over existing methods. This demonstrates that LLM agents can serve as an effective and scalable approach for creating state‑of‑the‑art MIAs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

On Reliability of Membership Inference Vulnerability Evaluation

The paper examines the reliability of membership inference attack (MIA) vulnerability evaluation. It identifies two weaknesses: finite‑sample bias from sampling shadow datasets from a fixed superset, and miscalibration when aggregating true positive rates across individuals at very low false positive rates. The authors propose simple fixes that avoid extra computational cost and suggest further improvements with additional computation.

By Joonas J\"alk\"o, Gauri Pradhan, Ossi R\"ais\"a, Antti Honkela
arXiv AI
2d ago

UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

UniGuardian is a training‑free detector for large language models that jointly identifies prompt injection, backdoor, and adversarial attacks—collectively called Prompt Trigger Attacks (PTA). It measures how structured prompt perturbations shift the model’s output distribution and uses a single‑forward strategy to detect attacks while generating text in a shared batched forward pass. Experiments show that UniGuardian accurately and efficiently identifies trigger‑activated prompts in LLMs.

By Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao