arXiv:2602. 05833v2 Announce Type: replace Abstract: There is a need for synthetic training and test datasets that replicate statistical distributions of original datasets without compromising their confidentiality.
By Laura Plein, Alexi Turcotte, Arina Hallemans, Andreas Zeller
arXiv:2603. 10937v2 Announce Type: replace Abstract: The use of synthetic data has become increasingly popular as a privacy-preserving alternative to sharing real datasets, especially in sensitive domains such as healthcare, finance, and demography.
By Rajdeep Pathak, Amit Basak, Sayantee Jana
arXiv:2607. 13541v1 Announce Type: cross Abstract: To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT).
By Na Li, Boyu Kuang, Hongsheng Hu, Liquan Chen, Hyoungshick Kim, Yansong Gao, Anmin Fu
arXiv:2405. 16361v4 Announce Type: replace Abstract: To protect privacy in regulated domains such as healthcare and finance, model owners may allow only remote API access while keeping both the training data and model parameters private.
By Kexin Li, Aastha Mehta, David Lie
arXiv:2606. 11267v1 Announce Type: new Abstract: Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise.
By Laurence A. Jacobs
arXiv:2606. 17110v1 Announce Type: cross Abstract: Large Language Models are increasingly trained on proprietary or sensitive data, from private healthcare and financial records to user conversations containing secrets.
By Md Abdullah Al Mamun, Ngoc Phu Doan, Pedram Zaree, Ihsen Alouani, Nael Abu-Ghazaleh
arXiv:2503. 23536v3 Announce Type: replace-cross Abstract: Unlearnable data (ULD) has emerged as an innovative defense technique to prevent machine learning models from learning meaningful patterns from specific data, thus protecting data privacy and security.
By Jiahao Li, Yiqiang Chen, Yunbing Xing, Yang Gu, Xiangyuan Lan
The paper investigates whether inexpensive spectral metrics from the heavy‑tailed self‑regularisation framework can predict membership inference attack (MIA) vulnerability, offering a scalable alternative to costly shadow‑model attacks. Experiments on image and tabular classification tasks show that stable rank correlates positively with overall MIA success, while Log alpha‑Norm correlates negatively with MIA risk in low false‑positive regimes, outperforming conventional generalisation gap measures. These findings suggest that neural network spectra contain privacy leakage signals not captured by traditional overfitting metrics, pointing to spectral analysis as a promising direction for privacy auditing.
By Richard J. Preen, Jim Smith
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
arXiv:2607. 12354v1 Announce Type: new Abstract: In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks.
By Rasmus Torp, Shailen K. Smith, Adam Breuer
The paper compares four audit methods for assessing identity‑level differential privacy in pre‑trained, black‑box face generators. Each method—GaussMech, KDE‑LR, MMD‑TV, and ROC‑HT—has distinct assumptions, hyperparameters, and finite‑sample limitations, and they produce markedly different epsilon estimates when applied to FaceFusion and InstantID. The study finds that all methods reveal significant identity distinguishability, but none can be reliably ranked in this high‑distinguishability regime, suggesting that future work should evaluate them on partially private mechanisms.
By Arman Zareian Jahromi, Vishnu Bondalakunta, Mohammad Akbar Bin Shah, Naimul Haque, Shuangqing Wei, George T. Amariucai
arXiv:2507. 01752v4 Announce Type: replace-cross Abstract: Gradient-based optimization is the workhorse of deep learning, offering efficient and scalable training via backpropagation.
By Ismail Labiad, Mathurin Videau, Matthieu Kowalski, Marc Schoenauer, Alessandro Leite, Julia Kempe, Olivier Teytaud