arXiv Machine Learning

PRIVET: PRoximIty leakage detection Via Extreme value Theory

arXiv Machine Learning
Jul 16

When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training

arXiv:2607. 13541v1 Announce Type: cross Abstract: To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT).

By Na Li, Boyu Kuang, Hongsheng Hu, Liquan Chen, Hyoungshick Kim, Yansong Gao, Anmin Fu
arXiv AI
Jul 21

A Survey on Unlearnable Data

arXiv:2503. 23536v3 Announce Type: replace-cross Abstract: Unlearnable data (ULD) has emerged as an innovative defense technique to prevent machine learning models from learning meaningful patterns from specific data, thus protecting data privacy and security.

By Jiahao Li, Yiqiang Chen, Yunbing Xing, Yang Gu, Xiangyuan Lan
arXiv Machine Learning
Sep 11

Predicting Privacy Leakage from Weight Spectral Density

The paper investigates whether inexpensive spectral metrics from the heavy‑tailed self‑regularisation framework can predict membership inference attack (MIA) vulnerability, offering a scalable alternative to costly shadow‑model attacks. Experiments on image and tabular classification tasks show that stable rank correlates positively with overall MIA success, while Log alpha‑Norm correlates negatively with MIA risk in low false‑positive regimes, outperforming conventional generalisation gap measures. These findings suggest that neural network spectra contain privacy leakage signals not captured by traditional overfitting metrics, pointing to spectral analysis as a promising direction for privacy auditing.

By Richard J. Preen, Jim Smith
arXiv Machine Learning
Aug 19

Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images

The paper compares four audit methods for assessing identity‑level differential privacy in pre‑trained, black‑box face generators. Each method—GaussMech, KDE‑LR, MMD‑TV, and ROC‑HT—has distinct assumptions, hyperparameters, and finite‑sample limitations, and they produce markedly different epsilon estimates when applied to FaceFusion and InstantID. The study finds that all methods reveal significant identity distinguishability, but none can be reliably ranked in this high‑distinguishability regime, suggesting that future work should evaluate them on partially private mechanisms.

By Arman Zareian Jahromi, Vishnu Bondalakunta, Mohammad Akbar Bin Shah, Naimul Haque, Shuangqing Wei, George T. Amariucai