Hugging Face Trending Papers

Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data

Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups.

arXiv Machine Learning
1d ago

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.

By Georgi Ganev, Emiliano De Cristofaro
arXiv Machine Learning
Sep 11

PEARL: A Task-Aware Framework for Evaluating Differentially Private Synthetic Educational Data

PEARL is a task-aware framework that evaluates differentially private synthetic educational data by checking validity, privacy protection, predictive usefulness, and suitability for the intended educational task. In a study of 96 settings, only 12 datasets passed all PEARL checks, with many failures due to missing outcome groups or distorted learning activity order. Even datasets that met privacy and predictive-usefulness criteria sometimes exhibited fairness issues and failed to support knowledge-tracing models, indicating that privacy alone does not guarantee practical usefulness.

By Xianghui Meng, Yujing Zhang, Jionghao Lin
Hugging Face Trending Papers
Jun 11

Disparate Impact in Synthetic Data Generation

We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data.

arXiv Machine Learning
Sep 15

Disparate Impact in Synthetic Data Generation

arXiv:2606.13105v2 Announce Type: replace Abstract: We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is t...

By Paul Andrey, Micha\"el Perrot, Batiste Le Bars, Marc Tommasi