arXiv AI

Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data

arXiv:2606. 16952v1 Announce Type: cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.

arXiv Machine Learning
1d ago

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.

By Georgi Ganev, Emiliano De Cristofaro
arXiv Machine Learning
Jun 2

Causal Evaluation of Membership Inference Attacks

arXiv:2602. 02819v4 Announce Type: replace Abstract: Membership Inference Attacks (MIAs) aim to distinguish training points (members) from unseen data (non-members), and are widely used to quantify memorization and assess privacy risks.

By Mathieu Even, Cl\'ement Berenfeld, Linus Bleistein, Tudor Cebere, Julie Josse, Aur\'elien Bellet
arXiv Machine Learning
Sep 17

QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing

QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.

By Shuaiqi Wang, Zinan Lin, Giulia Fanti