Disparate Impact in Synthetic Data Generation
arXiv:2606.13105v2 Announce Type: replace Abstract: We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is t...
We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data.
arXiv:2606.13105v2 Announce Type: replace Abstract: We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is t...
arXiv:2607. 07471v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness.
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups.
arXiv:2507. 19700v2 Announce Type: replace Abstract: We propose a new framework for generating tabular synthetic datasets via disjoint generative models.
arXiv:2604. 07486v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation.
arXiv:2602. 05833v2 Announce Type: replace Abstract: There is a need for synthetic training and test datasets that replicate statistical distributions of original datasets without compromising their confidentiality.
arXiv:2606. 08259v1 Announce Type: new Abstract: This paper investigates the problem of generating synthetic tabular data with differential privacy (DP) guarantees, enabling data sharing in sensitive domains.
The paper investigates how to maintain causal fairness when releasing synthetic data by applying the DECAF framework across nine different synthetic data generators from three families (marginals-based, GAN, and diffusion) and three levels of differential privacy. Experiments on Adult and COMPAS datasets show that the causal diffusion backbone consistently produces the fairest data with fidelity comparable to the marginals tier, while the fairness cuts have minimal impact on downstream classifier performance and do not degrade privacy guarantees.
The paper demonstrates that causal fairness mechanisms can be applied across a wide range of synthetic data generators, including marginal‑based, GAN, and diffusion models, each with differentially private variants. By porting three fairness definitions to nine generators and testing them on Adult and COMPAS datasets, the authors show that the causal diffusion backbone consistently produces the fairest data releases while maintaining high fidelity. The fairness cuts have minimal impact on data quality, costing downstream classifiers only about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees does not reduce fairness.
Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-c...
arXiv:2609.36678v1 Announce Type: new Abstract: Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm,...
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.