Hugging Face Trending Papers

Disparate Impact in Synthetic Data Generation

We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data.

arXiv Machine Learning
Sep 15

Disparate Impact in Synthetic Data Generation

arXiv:2606.13105v2 Announce Type: replace Abstract: We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is t...

By Paul Andrey, Micha\"el Perrot, Batiste Le Bars, Marc Tommasi
arXiv Machine Learning
Jun 9

Disjoint Generation of Synthetic Data

arXiv:2507. 19700v2 Announce Type: replace Abstract: We propose a new framework for generating tabular synthetic datasets via disjoint generative models.

By Anton Danholt Lautrup, Muhammad Rajabinasab, Tobias Hyrup, Arthur Zimek, Peter Schneider-Kamp
Hugging Face Trending Papers
Sep 2

Portable Causal Fairness Across Synthetic Data Generator Families

The paper investigates how to maintain causal fairness when releasing synthetic data by applying the DECAF framework across nine different synthetic data generators from three families (marginals-based, GAN, and diffusion) and three levels of differential privacy. Experiments on Adult and COMPAS datasets show that the causal diffusion backbone consistently produces the fairest data with fidelity comparable to the marginals tier, while the fairness cuts have minimal impact on downstream classifier performance and do not degrade privacy guarantees.

arXiv Machine Learning
Sep 4

Portable Causal Fairness Across Synthetic Data Generator Families

The paper demonstrates that causal fairness mechanisms can be applied across a wide range of synthetic data generators, including marginal‑based, GAN, and diffusion models, each with differentially private variants. By porting three fairness definitions to nine generators and testing them on Adult and COMPAS datasets, the authors show that the causal diffusion backbone consistently produces the fairest data releases while maintaining high fidelity. The fairness cuts have minimal impact on data quality, costing downstream classifiers only about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees does not reduce fairness.

By Steven Golob, Sikha Pentyala, Martine De Cock